The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A review of 160 studies shows NLP has scaled analysis of sustainability reports but remains fragmented: lexicons, assisted coding and topic models dominate while linguistic variation, temporal drift and methodological heterogeneity make cross‑study comparisons unreliable, so the field needs shared benchmarks, gold‑label datasets and replication standards.

From Dictionaries to Deep Learning: A Systematic Mapping Review of the Natural Language Processing Tasks Used to Analyze Sustainability Reports
Hannes Cordes · August 31, 2026 · Corporate Social Responsibility and Environmental Management
openalex review_meta n/a evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Hannes Cordes provider ID

Semantic Scholar

Latest observation:

  1. Hannes Cordes unresolved corpus identity
A systematic review of 160 studies finds that keyword/lexicon searches, computer‑assisted content analysis, and topic modeling are the most common NLP approaches to sustainability reporting, but heterogeneous methods, linguistic/cultural variation, and temporal drift undermine comparability and introduce bias, prompting calls for benchmarks, shared datasets, and replication.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

ABSTRACT This study aims at shedding light on the vast landscape of natural language processing (NLP) tasks used when analyzing sustainability reports or sustainability within integrated annual reports. A systematic literature review is carried out, identifying 160 studies of relevance. The results are presented as a hierarchical categorization, grouping methods and models into tasks, and tasks into categories. Most common tasks are keyword searches, (computer aided) content analysis, and topic modeling. Researchers apply NLP most frequently in a comparative setting, investigating relations or exploring changes over time. Common limitations are different linguistic and cultural features, temporal changes in reporting language, and the risk to introduce or amplify bias. This study provides researchers and practitioners with guidance and recommendations when applying NLP tasks. It additionally proposes a research agenda focused on guidelines and benchmarks as well as replication and verification studies.

Summary

Main Finding

A systematic literature review of 160 studies maps the landscape of NLP methods applied to sustainability reporting (including integrated annual reports). The authors produce a hierarchical categorization of methods → tasks → task categories and find that the most common NLP tasks are keyword searches, (computer-aided) content analysis, and topic modeling. NLP is most frequently applied in comparative designs (cross‑sectional or longitudinal) to investigate relationships or changes over time. Key limitations are linguistic/cultural variation, temporal drift in reporting language, and the potential to introduce or amplify bias. The paper offers guidance and a research agenda emphasizing guidelines, benchmarks, and replication/verification studies.

Key Points

  • Scope: Systematic review covering 160 relevant studies on NLP applications to sustainability / integrated reporting.
  • Taxonomy: Methods and models are grouped into tasks, and tasks into higher‑level categories (hierarchical organization for researchers/practitioners).
  • Most common tasks:
    • Keyword searches (lexicon-based measures)
    • Computer-aided content analysis (manual-assisted coding workflows)
    • Topic modeling (unsupervised discovery of themes)
  • Typical uses: Comparative analyses (cross-firm, cross-country) and temporal studies tracking language or disclosure changes.
  • Common limitations and risks:
    • Linguistic and cultural heterogeneity across reports and jurisdictions
    • Temporal changes (drift) in reporting language and terminology
    • Risk of bias introduction or amplification from model choices, training data, or lexicons
  • Outputs: Practical guidance and recommendations for applying NLP in this domain plus a proposed research agenda focused on standards and reproducibility.

Data & Methods

  • Methodology: Systematic literature review (selection and synthesis of peer-reviewed studies and relevant work).
  • Sample: 160 studies identified and analyzed.
  • Analytical approach: Development of a hierarchical categorization that maps individual methods/models into discrete NLP tasks, and groups tasks into broader categories — enabling structured comparison of what techniques are used and how.
  • Evidence base: Aggregated descriptions of task frequencies, application settings (comparative, longitudinal), reported limitations, and methodological recommendations across the reviewed literature.

Implications for AI Economics

  • Measurement validity and comparability:
    • Heterogeneous NLP approaches (lexicons, supervised models, unsupervised topic models) produce different indicators of sustainability disclosure; lack of standards hampers cross-study comparability and meta-analysis.
    • Temporal drift and multilingual reporting create measurement error—important for empirical work linking disclosure language to economic outcomes (e.g., cost of capital, firm valuation, regulatory responses).
  • Bias and inference risks:
    • Unchecked biases in lexicons or training data can distort estimated relationships between textual measures and economic variables, leading to misleading policy or managerial conclusions.
    • Causal inference using NLP-derived variables requires careful validation (e.g., robustness checks, alternative text measures, instrumenting where appropriate).
  • Research and infrastructure needs:
    • Standardized benchmarks, shared gold‑label datasets, and agreed evaluation metrics for sustainability-report NLP would improve reproducibility and comparability.
    • Replication and verification studies are necessary to vet widely used measures before they are used in high‑stakes economic or regulatory analysis.
    • Time-aware and cross‑lingual models (or domain adaptation techniques) are needed to handle reporting drift and cultural/linguistic heterogeneity.
  • Practical recommendations for researchers and practitioners:
    • Share code, lexicons, and labeled datasets; document preprocessing and model choices.
    • Use human-in-the-loop validation and report uncertainty or sensitivity to lexicon/model choices.
    • Prefer validated benchmarks when available; otherwise conduct robustness checks across multiple NLP methods.
    • Consider multilingual pipelines and temporal calibration (retraining or domain adaptation) to reduce measurement bias.
  • Policy and market implications:
    • Standardization of disclosure language or reporting formats (or official benchmarks) would reduce measurement frictions for researchers and investors.
    • Regulators and standard-setters should account for NLP measurement limits when using automated approaches for monitoring or enforcement.

Overall, the review highlights both the promise of NLP for scaling sustainability-report analysis and the urgent methodological tasks—standards, benchmarks, replication, and bias mitigation—that AI economics researchers must address to ensure reliable empirical inference.

Assessment

Paper Typereview_meta Evidence Strengthn/a — This is a systematic literature review and taxonomy rather than an original causal or empirical identification study, so it does not provide primary causal evidence to rate. Methods Rigorhigh — The paper applies a systematic review protocol across 160 studies, constructs a clear hierarchical taxonomy of methods→tasks→task categories, and synthesizes recurring applications, limitations, and recommendations; these features indicate careful, reproducible synthesis though details on search strategy, inclusion criteria, and coding reliability (not provided here) would determine the top score. SampleA curated set of 160 peer-reviewed and relevant studies applying NLP to sustainability reporting and integrated annual reports; covers a range of methods (lexicon/keyword searches, computer-aided content analysis, topic models, supervised models), study designs (cross-sectional comparative, longitudinal temporal analyses), and reporting jurisdictions/languages. Themesgovernance adoption GeneralizabilityFindings summarize the applied-NLP literature on sustainability/integrated reporting and may not generalize to other corporate text domains (earnings calls, analyst reports, internal documents)., Likely biased toward studies published in dominant languages (e.g., English) and common jurisdictions; cross-lingual and less-studied contexts are underrepresented., Temporal coverage means older methodological patterns (e.g., lexicons, topic models) may dominate even as newer models (large language models, transformer fine-tuning) gain use—so frequency counts may lag methodological change., Conclusions on best practices reflect aggregate reporting and author recommendations rather than validated, universally applicable benchmarks., Quality depends on the completeness of the literature search and coding reliability; publication bias (positive/novel-method studies) may affect patterns observed.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The systematic literature review identified and analyzed 160 studies on NLP applications to sustainability reporting, including integrated annual reports. Other positive Number of relevant studies included in the review
Reading fidelity high
Study strength medium
n=160
160 studies
0.24
The review develops a hierarchical categorization that maps NLP methods and models into tasks and groups tasks into broader task categories. Other positive Structured classification of NLP methods and tasks
Reading fidelity high
Study strength medium
n=160
0.24
Keyword searches, computer-aided content analysis, and topic modeling are the most common NLP tasks used in the reviewed sustainability-reporting literature. Adoption Rate positive Frequency of NLP task use across reviewed studies
Reading fidelity high
Study strength medium
n=160
0.24
NLP is most frequently applied in comparative research designs, including cross-sectional and longitudinal studies, to investigate relationships or changes over time. Other positive Prevalence of comparative, cross-sectional, and longitudinal study designs
Reading fidelity high
Study strength medium
n=160
0.24
Linguistic and cultural heterogeneity across reports and jurisdictions is a key limitation for NLP-based sustainability-report analysis. Error Rate negative Validity and comparability of NLP-derived sustainability-report measures
Reading fidelity high
Study strength medium
n=160
0.24
Temporal drift in reporting language and terminology is a key limitation for applying NLP methods consistently over time. Error Rate negative Temporal comparability and measurement validity of NLP-derived text indicators
Reading fidelity high
Study strength medium
n=160
0.24
NLP model choices, training data, and lexicons can introduce or amplify bias in sustainability-report measures. Ai Safety And Ethics negative Bias and reliability of NLP-derived sustainability-report indicators
Reading fidelity high
Study strength medium
n=160
0.24
The review recommends standardized benchmarks, shared gold-label datasets, agreed evaluation metrics, and replication or verification studies to improve reproducibility and comparability. Research Productivity positive Reproducibility and cross-study comparability of NLP-based sustainability-report research
Reading fidelity high
Study strength medium
n=160
0.24
The review recommends sharing code, lexicons, and labeled datasets, documenting preprocessing and model choices, and validating NLP measures with human-in-the-loop procedures and robustness checks. Research Productivity positive Transparency, validation, and robustness of NLP-based measurement
Reading fidelity high
Study strength medium
n=160
0.24

Notes