The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI sharpens the lens on ESG reporting—transformer-based NLP scales comparability and flags potential greenwashing—but evidence that managerial deployment of AI actually boosts stakeholder trust or market outcomes is scarce, and language and explainability shortfalls constrain practical use.

The Role of Artificial Intelligence in Enhancing ESG Disclosure Quality in Accounting
Jiacheng Liu, Ye Yuan, Zhelun Zhu · January 09, 2026 · Journal of risk and financial management
openalex review_meta medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Jiacheng Liu exact ORCID
  2. Ye Yuan provider ID
  3. Zhelun Zhu provider ID

Semantic Scholar

Latest observation:

  1. Jiacheng Liu provider ID
  2. Ye Yuan provider ID
  3. Zhelun Zhu provider ID
AI and modern NLP substantially improve scalable measurement, comparability, and preliminary detection of greenwashing in ESG disclosures, but causal evidence that firms’ managerial adoption of AI meaningfully improves stakeholder outcomes is limited and hindered by multilingual biases and interpretability gaps.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

As corporate sustainability reporting evolves into a pivotal resource for investors, regulators, and stakeholders, the imperative to evaluate and elevate ESG disclosure quality intensifies amid persistent challenges like opacity, inconsistency, and greenwashing. This review synthesizes interdisciplinary insights from accounting, finance, and computational linguistics on artificial intelligence (AI), particularly natural language processing (NLP) and machine learning (ML), as a transformative force in this domain. We delineate ESG disclosure quality across four operational dimensions: readability, comparability, informativeness, and credibility. By integrating cutting-edge methodological innovations (e.g., transformer-based models for semantic analysis), empirical linkages between AI-extracted signals and market/governance outcomes, and normative discussions on AI’s auditing potential, we demonstrate AI’s efficacy in scaling measurement, harmonizing heterogeneous narratives, and prototyping greenwashing detection. Nonetheless, causal evidence linking managerial AI adoption to stakeholder-perceived enhancements remains limited, compounded by biases in multilingual applications and interpretability deficits. We propose a forward-looking agenda, prioritizing cross-lingual benchmarking, curated greenwashing datasets, AI-assurance pilots, and interpretability standards, to harness AI for substantive, equitable improvements in ESG reporting and accountability.

Summary

Main Finding

AI—especially modern NLP/ML (transformers, embeddings)—is a powerful, scalable tool for measuring and improving ESG disclosure quality across readability, comparability, informativeness, and credibility. Empirical work shows AI-extracted ESG signals map to market and governance outcomes and can prototype greenwashing detection, but causal evidence that managerial AI adoption improves stakeholder-perceived disclosure is limited. Important gaps remain in multilingual fairness, interpretability, benchmark datasets, and operationalized AI assurance.

Key Points

  • Four operational dimensions of ESG disclosure quality:
    • Readability: linguistic complexity, clarity, and organization of reports.
    • Comparability: capacity to harmonize heterogeneous narratives across firms/sectors/standards.
    • Informativeness: incremental information content relevant to valuation, risk, and stewardship.
    • Credibility: truthfulness and resistance to greenwashing or selective disclosure.
  • AI strengths:
    • Scalability: automates extraction across large corpora, longitudinal coverage, and frequent updates.
    • Semantic richness: transformer models capture nuanced meaning and context beyond bag-of-words.
    • Harmonization: embeddings and supervised classifiers align heterogeneous disclosures to common taxonomies.
    • Detection prototypes: supervised and unsupervised methods flag likely greenwashing, omission, or rhetorical obfuscation.
  • Empirical linkages:
    • AI-derived ESG signals are associated with stock-price reactions, analyst behavior, and governance actions in several studies.
    • Evidence supports AI improving cross-firm comparability and revealing hidden patterns not visible in structured filings.
  • Limitations and risks:
    • Causality: few studies identify causal effects of firms’ managerial adoption of AI on disclosure quality or stakeholder trust.
    • Multilingual and cross-jurisdictional bias: models trained on dominant-language corpora perform worse for other languages/regions.
    • Interpretability and auditability: black-box models complicate assurance, regulatory use, and stakeholder trust.
    • Data limitations: sparse, inconsistent, or proprietary labels (e.g., greenwashing) hinder supervised learning and benchmarking.
  • Normative debate:
    • AI offers potential to augment auditing and regulatory surveillance, but raises questions on standards, oversight, and accountability for automated assessments.

Data & Methods

  • Data sources commonly used:
    • Corporate sustainability reports, annual reports (MD&A, ESG sections), proxy statements, regulatory filings, press releases, social media, third-party ESG ratings, and satellite or supply-chain data for verification.
    • Annotated corpora are limited; many studies rely on distant supervision (e.g., mapping to existing ratings) or manual annotation on small samples.
  • NLP/ML methods:
    • Traditional: dictionary approaches (LIWC-style), bag-of-words, TF-IDF, topic models (LDA).
    • Modern: contextualized embeddings (BERT and variants), transformer-based classifiers (fine-tuned for ESG categories), sentence/document embeddings for clustering and similarity, cross-lingual transformers for multilingual tasks.
    • Supervised learning: binary/multiclass classifiers for greenwashing, issue-level disclosure, sentiment, and forward-looking statements.
    • Unsupervised/semi-supervised: anomaly detection for inconsistent narratives, clustering to discover disclosure patterns.
    • Explainability tools: feature importance, attention analysis, and post-hoc methods (SHAP/LIME) used but remain imperfect for rigorous assurance.
  • Empirical strategies:
    • Correlational analyses linking AI-extracted features to stock returns, trading volume, analyst forecasts, governance events.
    • Panel regressions controlling for firm-level covariates; some event-study designs around report releases.
    • Few causal identification strategies (e.g., difference-in-differences tied to regulatory changes or exogenous IT shocks); these are underrepresented.
  • Methodological gaps highlighted:
    • Need for cross-lingual, domain-adaptive benchmarks.
    • Insufficient curated, labeled datasets for greenwashing and credibility.
    • Limited use of interpretable models or methods designed for auditability.

Implications for AI Economics

  • Information asymmetry & market efficiency:
    • AI-derived ESG signals can reduce search costs and information frictions, potentially improving price discovery for ESG-related risks and opportunities.
    • Better comparability may compress mispricing premia tied to disclosure opacity, affecting returns to disclosure and quality.
  • Corporate incentives and competition:
    • Scalability of ESG measurement raises stakes: firms may face stronger market/regulatory scrutiny, altering their incentives toward substantive disclosure or sophisticated rhetoric (and therefore an arms race in detection).
    • Demand for AI tools and expertise creates returns to scale for incumbents (consultancies, data providers) and could concentrate informational advantages.
  • Policy and regulation:
    • Regulators can leverage AI for surveillance and enforcement but must set standards for model transparency, cross-jurisdictional fairness, and evidentiary thresholds for findings (to avoid false positives/negative consequences).
    • AI-assurance pilots (algorithmic audit trails, hybrid human-AI review) could operationalize mandatory ESG assurance and reduce audit costs, but require standardization.
  • Labor and service markets:
    • Growth in AI tools for ESG analytics may shift demand toward data annotation, model validation, and interpretability specialists; could reduce routine disclosure assessment roles while increasing higher-skill assurance work.
  • Research agenda implications for economics:
    • Causal evidence: economists can design randomized or quasi-experimental studies assessing managerial AI adoption, disclosure responses, and investor reactions.
    • Welfare analysis: quantify social value of improved ESG transparency (e.g., externality internalization, capital allocation efficiency) and potential distributional effects across firms and countries.
    • Measurement and bias correction: develop econometric techniques to adjust AI-extracted signals for measurement error and language/cultural biases to enable valid cross-country comparisons.
  • Practical priorities to realize benefits:
    • Develop cross-lingual benchmarking suites and open annotated greenwashing datasets.
    • Fund and pilot AI-assurance frameworks that combine interpretable models, human oversight, and regulatory engagement.
    • Establish interpretability and documentation standards (model cards, dataset sheets) for ESG-AI tools to support accountability and economic research replication.

Assessment

Paper Typereview_meta Evidence Strengthmedium — The paper synthesizes multiple empirical studies that show correlations between AI/NLP-extracted ESG signals and market or governance outcomes and demonstrates strong methodological progress (e.g., transformer models). However, causal identification is weak: few randomized or quasi-experimental designs, limited causal evidence that managerial AI adoption improves stakeholder-perceived outcomes, and many results are observational and context-specific. Methods Rigormedium — The review integrates state-of-the-art computational methods (transformers, semantic analysis) and cross-disciplinary empirical findings, indicating careful methodological coverage; nevertheless, rigor is constrained by heterogeneity in primary studies (varying datasets, labels, and evaluation metrics), uneven benchmarking, and acknowledged gaps in multilingual validation and interpretability. SampleA systematic, interdisciplinary literature synthesis drawing on accounting, finance, and computational linguistics studies; primary empirical work reviewed typically uses corpora of corporate ESG reports, regulatory filings, and sustainability disclosures (often English-heavy), transformer- and ML-based semantic models, curated greenwashing/proxy labels in limited datasets, and firm-level market/governance outcome data (stock returns, analyst reactions, governance indices); also includes conceptual pieces and pilot case studies of AI-assisted assurance. Themesgovernance adoption human_ai_collab GeneralizabilityEnglish-language dataset bias limits cross-lingual applicability, Heterogeneous ESG standards and reporting formats reduce comparability across jurisdictions, Empirical work concentrated on large, publicly listed firms and certain sectors, Many findings are correlational and may not generalize causally across settings, Rapid evolution of AI models (model/version drift) may change performance over time, Regulatory and institutional differences across countries constrain external validity

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
ESG disclosure quality can be delineated across four operational dimensions: readability, comparability, informativeness, and credibility. Output Quality positive ESG disclosure quality (readability, comparability, informativeness, credibility)
Reading fidelity high
Study strength speculative
not reported
0.04
AI (particularly NLP and ML, including transformer-based models) is a transformative force in ESG disclosure, enabling scaling of measurement, harmonizing heterogeneous narratives, and prototyping greenwashing detection. Output Quality positive ability to measure and harmonize ESG disclosures and detect greenwashing
Reading fidelity high
Study strength medium
not reported
0.24
There exist empirical linkages between AI-extracted signals from ESG disclosures and market and governance outcomes. Governance And Regulation mixed market and governance outcomes associated with AI-extracted ESG signals
Reading fidelity high
Study strength medium
not reported
0.24
Causal evidence linking managerial AI adoption to stakeholder-perceived enhancements in ESG reporting is limited. Output Quality negative causal impact of managerial AI adoption on stakeholder perceptions of ESG reporting
Reading fidelity high
Study strength high
not reported
0.4
Multilingual applications of AI for ESG disclosure suffer from biases and there are interpretability deficits in current AI approaches. Ai Safety And Ethics negative bias and interpretability limitations in AI models applied to ESG disclosures
Reading fidelity high
Study strength medium
not reported
0.24
AI has potential as an auditing tool (AI’s auditing potential) to support assurance of ESG disclosures. Regulatory Compliance positive potential of AI to support auditing/assurance of ESG disclosures
Reading fidelity high
Study strength speculative
not reported
0.04
The paper proposes a forward-looking research and practice agenda prioritizing cross-lingual benchmarking, curated greenwashing datasets, AI-assurance pilots, and interpretability standards to improve ESG reporting and accountability. Governance And Regulation positive research and practice priorities for improving ESG reporting via AI
Reading fidelity high
Study strength speculative
not reported
0.04
Transformer-based models enable improved semantic analysis of ESG narratives, facilitating harmonization of heterogeneous textual disclosures. Output Quality positive semantic analysis quality and harmonization of ESG textual disclosures
Reading fidelity high
Study strength medium
not reported
0.24

Notes