0 cumulative citations
View corpus contextA review of 160 studies shows NLP has scaled analysis of sustainability reports but remains fragmented: lexicons, assisted coding and topic models dominate while linguistic variation, temporal drift and methodological heterogeneity make cross‑study comparisons unreliable, so the field needs shared benchmarks, gold‑label datasets and replication standards.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextABSTRACT This study aims at shedding light on the vast landscape of natural language processing (NLP) tasks used when analyzing sustainability reports or sustainability within integrated annual reports. A systematic literature review is carried out, identifying 160 studies of relevance. The results are presented as a hierarchical categorization, grouping methods and models into tasks, and tasks into categories. Most common tasks are keyword searches, (computer aided) content analysis, and topic modeling. Researchers apply NLP most frequently in a comparative setting, investigating relations or exploring changes over time. Common limitations are different linguistic and cultural features, temporal changes in reporting language, and the risk to introduce or amplify bias. This study provides researchers and practitioners with guidance and recommendations when applying NLP tasks. It additionally proposes a research agenda focused on guidelines and benchmarks as well as replication and verification studies.
Summary
Main Finding
A systematic literature review of 160 studies maps the landscape of NLP methods applied to sustainability reporting (including integrated annual reports). The authors produce a hierarchical categorization of methods → tasks → task categories and find that the most common NLP tasks are keyword searches, (computer-aided) content analysis, and topic modeling. NLP is most frequently applied in comparative designs (cross‑sectional or longitudinal) to investigate relationships or changes over time. Key limitations are linguistic/cultural variation, temporal drift in reporting language, and the potential to introduce or amplify bias. The paper offers guidance and a research agenda emphasizing guidelines, benchmarks, and replication/verification studies.
Key Points
- Scope: Systematic review covering 160 relevant studies on NLP applications to sustainability / integrated reporting.
- Taxonomy: Methods and models are grouped into tasks, and tasks into higher‑level categories (hierarchical organization for researchers/practitioners).
- Most common tasks:
- Keyword searches (lexicon-based measures)
- Computer-aided content analysis (manual-assisted coding workflows)
- Topic modeling (unsupervised discovery of themes)
- Typical uses: Comparative analyses (cross-firm, cross-country) and temporal studies tracking language or disclosure changes.
- Common limitations and risks:
- Linguistic and cultural heterogeneity across reports and jurisdictions
- Temporal changes (drift) in reporting language and terminology
- Risk of bias introduction or amplification from model choices, training data, or lexicons
- Outputs: Practical guidance and recommendations for applying NLP in this domain plus a proposed research agenda focused on standards and reproducibility.
Data & Methods
- Methodology: Systematic literature review (selection and synthesis of peer-reviewed studies and relevant work).
- Sample: 160 studies identified and analyzed.
- Analytical approach: Development of a hierarchical categorization that maps individual methods/models into discrete NLP tasks, and groups tasks into broader categories — enabling structured comparison of what techniques are used and how.
- Evidence base: Aggregated descriptions of task frequencies, application settings (comparative, longitudinal), reported limitations, and methodological recommendations across the reviewed literature.
Implications for AI Economics
- Measurement validity and comparability:
- Heterogeneous NLP approaches (lexicons, supervised models, unsupervised topic models) produce different indicators of sustainability disclosure; lack of standards hampers cross-study comparability and meta-analysis.
- Temporal drift and multilingual reporting create measurement error—important for empirical work linking disclosure language to economic outcomes (e.g., cost of capital, firm valuation, regulatory responses).
- Bias and inference risks:
- Unchecked biases in lexicons or training data can distort estimated relationships between textual measures and economic variables, leading to misleading policy or managerial conclusions.
- Causal inference using NLP-derived variables requires careful validation (e.g., robustness checks, alternative text measures, instrumenting where appropriate).
- Research and infrastructure needs:
- Standardized benchmarks, shared gold‑label datasets, and agreed evaluation metrics for sustainability-report NLP would improve reproducibility and comparability.
- Replication and verification studies are necessary to vet widely used measures before they are used in high‑stakes economic or regulatory analysis.
- Time-aware and cross‑lingual models (or domain adaptation techniques) are needed to handle reporting drift and cultural/linguistic heterogeneity.
- Practical recommendations for researchers and practitioners:
- Share code, lexicons, and labeled datasets; document preprocessing and model choices.
- Use human-in-the-loop validation and report uncertainty or sensitivity to lexicon/model choices.
- Prefer validated benchmarks when available; otherwise conduct robustness checks across multiple NLP methods.
- Consider multilingual pipelines and temporal calibration (retraining or domain adaptation) to reduce measurement bias.
- Policy and market implications:
- Standardization of disclosure language or reporting formats (or official benchmarks) would reduce measurement frictions for researchers and investors.
- Regulators and standard-setters should account for NLP measurement limits when using automated approaches for monitoring or enforcement.
Overall, the review highlights both the promise of NLP for scaling sustainability-report analysis and the urgent methodological tasks—standards, benchmarks, replication, and bias mitigation—that AI economics researchers must address to ensure reliable empirical inference.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The systematic literature review identified and analyzed 160 studies on NLP applications to sustainability reporting, including integrated annual reports. Other | positive | Number of relevant studies included in the review |
Reading fidelity
high
Study strength
medium
|
n=160
160 studies
|
| The review develops a hierarchical categorization that maps NLP methods and models into tasks and groups tasks into broader task categories. Other | positive | Structured classification of NLP methods and tasks |
Reading fidelity
high
Study strength
medium
|
n=160
|
| Keyword searches, computer-aided content analysis, and topic modeling are the most common NLP tasks used in the reviewed sustainability-reporting literature. Adoption Rate | positive | Frequency of NLP task use across reviewed studies |
Reading fidelity
high
Study strength
medium
|
n=160
|
| NLP is most frequently applied in comparative research designs, including cross-sectional and longitudinal studies, to investigate relationships or changes over time. Other | positive | Prevalence of comparative, cross-sectional, and longitudinal study designs |
Reading fidelity
high
Study strength
medium
|
n=160
|
| Linguistic and cultural heterogeneity across reports and jurisdictions is a key limitation for NLP-based sustainability-report analysis. Error Rate | negative | Validity and comparability of NLP-derived sustainability-report measures |
Reading fidelity
high
Study strength
medium
|
n=160
|
| Temporal drift in reporting language and terminology is a key limitation for applying NLP methods consistently over time. Error Rate | negative | Temporal comparability and measurement validity of NLP-derived text indicators |
Reading fidelity
high
Study strength
medium
|
n=160
|
| NLP model choices, training data, and lexicons can introduce or amplify bias in sustainability-report measures. Ai Safety And Ethics | negative | Bias and reliability of NLP-derived sustainability-report indicators |
Reading fidelity
high
Study strength
medium
|
n=160
|
| The review recommends standardized benchmarks, shared gold-label datasets, agreed evaluation metrics, and replication or verification studies to improve reproducibility and comparability. Research Productivity | positive | Reproducibility and cross-study comparability of NLP-based sustainability-report research |
Reading fidelity
high
Study strength
medium
|
n=160
|
| The review recommends sharing code, lexicons, and labeled datasets, documenting preprocessing and model choices, and validating NLP measures with human-in-the-loop procedures and robustness checks. Research Productivity | positive | Transparency, validation, and robustness of NLP-based measurement |
Reading fidelity
high
Study strength
medium
|
n=160
|