The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Leading multimodal AI systems struggle to detect greenwashing in Polish ESG reports, frequently flagging high-performing firms as more deceptive and showing virtually no agreement across models; current off‑the‑shelf tools conflate polished sustainability communication with fraudulent intent and lack the contextual grounding for reliable ESG assurance.

Evaluating Multimodal AI for Greenwashing Detection: A Comparative Analysis of ChatGPT, Claude, and Gemini in ESG Reports
Jacek Krzysztof Jakubczak, Dorota Chmielewska-Muciek, Katarzyna Iwanicka · December 25, 2025 · Sustainability
openalex correlational low evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Jacek Krzysztof Jakubczak provider ID
  2. Dorota Chmielewska-Muciek provider ID
  3. Katarzyna Iwanicka provider ID

Semantic Scholar

Latest observation:

  1. J. Jakubczak provider ID
  2. Dorota Chmielewska-Muciek provider ID
  3. Katarzyna Iwanicka provider ID
Three leading multimodal AI systems failed to reliably detect visual greenwashing in 20 Polish ESG reports—often assigning higher greenwashing scores to firms with stronger ESRS-aligned performance—and showed nearly zero agreement across tools.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The rapid expansion of sustainability reporting under the EU Corporate Sustainability Reporting Directive (CSRD) has intensified concerns about greenwashing, particularly in visual communication within ESG reports. Recent advances in multimodal artificial intelligence offer new possibilities for automated detection, yet their reliability in non-English corporate reporting contexts remains unclear. This study evaluates the greenwashing detection capabilities of three leading multimodal AI systems—ChatGPT 5.1, Claude 4.5 Sonnet, and Gemini 2.5 Flash—using a purposively selected sample of 20 Polish ESG reports benchmarked against ESRS-aligned performance scores from the national “Ranking ESG”. A standardized auditing prompt was applied across all tools to generate comparable assessments of visual greenwashing. Contrary to theoretical expectations and all four hypotheses, the models did not demonstrate negative correlations between performance and AI-detected greenwashing; instead, high-performing firms frequently received higher greenwashing scores. Dimensional analyses showed inconsistent and often contradictory evaluations across Environmental, Social, and Governance pillars, while inter-tool reliability proved extremely low (Krippendorff’s α ≈ 0). These findings indicate that current multimodal AI systems conflate communication sophistication with deceptive intent and lack sufficient contextual understanding for ESG assurance. The study highlights significant methodological limitations and outlines directions for developing domain-specific, ESRS-aligned AI tools for greenwashing detection.

Summary

Main Finding

Multimodal LLMs (ChatGPT 5.1, Claude 4.5 Sonnet, Gemini 2.5 Flash) failed to reliably detect visual greenwashing in non-English (Polish) ESG reports. Rather than showing the expected negative correlation between substantive ESG performance and AI-flagged greenwashing, the tools often flagged higher-performing firms more heavily. Across Environmental, Social, and Governance dimensions the models produced inconsistent and contradictory judgments, and inter-model agreement was effectively nil (Krippendorff’s α ≈ 0). Overall, current general-purpose multimodal systems conflate polished communication with deceptive intent and lack the contextual grounding required for trustworthy ESG assurance.

Key Points

  • Models evaluated: ChatGPT 5.1, Claude 4.5 Sonnet, Gemini 2.5 Flash.
  • Sample: purposively selected n = 20 Polish ESG reports, benchmarked to ESRS-aligned performance scores from the national “Ranking ESG”.
  • Procedure: a standardized auditing prompt applied identically across tools to assess visual greenwashing in reports.
  • Main unexpected result: AI greenwashing scores were not negatively correlated with ESRS-aligned performance; high-performing firms frequently received higher greenwashing scores.
  • Pillar-level analysis (E, S, G) produced inconsistent and often contradictory results across tools.
  • Inter-tool reliability was effectively zero (Krippendorff’s α ≈ 0), indicating no reproducible agreement among systems.
  • Interpretation: models appear to equate sophisticated visual/communicative quality with possible deception rather than assessing substantive performance or ESRS-conformant disclosures.
  • Study cautions against using off-the-shelf multimodal LLMs for ESG assurance in non-English contexts.

Data & Methods

  • Data: 20 Polish corporate ESG reports selected purposively to span a range of ESRS-aligned performance scores from the national “Ranking ESG”.
  • Benchmark: ESRS-aligned performance scores from Ranking ESG served as the substantive-performance benchmark.
  • Tools: three leading multimodal AI systems—ChatGPT 5.1, Claude 4.5 Sonnet, Gemini 2.5 Flash—accessed contemporaneously and prompted with the same standardized audit instruction.
  • Prompting: a single, standardized visual-audit prompt was used to generate comparable greenwashing assessments across tools; outputs were mapped to a common scoring rubric.
  • Analyses:
    • Correlation tests between AI greenwashing scores and ESRS-aligned performance scores (hypothesized negative correlation).
    • Dimensional breakdowns for Environmental, Social, and Governance pillars.
    • Reliability analysis across models using Krippendorff’s α.
  • Key methodological limitations noted by authors:
    • Small, purposive, single-country (Poland) sample limits generalizability.
    • Non-English reporting poses known challenges for models trained predominantly on English data.
    • A single standardized prompt may not capture optimal prompting strategies per model.
    • Off-the-shelf models not fine-tuned on ESRS taxonomy or domain-specific labeled datasets.

Implications for AI Economics

  • Market signals and capital allocation: If general-purpose multimodal AIs flag high-performing firms as greenwashing (false positives), automated screening tools could distort investor signals and lead to mispricing or misallocation of capital.
  • Demand for domain-specific assurance tools: Findings increase economic value for specialized, ESRS-aligned multimodal models or classifiers trained on multilingual, audit-labeled corpora—creating an opportunity for vendors and service providers in AI assurance markets.
  • Regulatory and compliance risk: Regulators and standard setters should be cautious about relying on off-the-shelf AI for ESG verification; poor reliability could produce regulatory errors and legal risks for firms and auditors.
  • Productivity and cost of assurance: While automated detection promises cost savings, current tools may raise verification costs (manual review of false positives/negatives) until domain-tuned systems reach acceptable accuracy and reliability.
  • Need for standardized benchmarks and public datasets: Economically efficient development of trustworthy tools requires labeled, multilingual, ESRS-aligned datasets and open benchmarks to reduce information asymmetries between AI vendors and users.
  • Incentives and strategic behavior: If firms learn that polished presentation triggers automated scrutiny, they may alter communication strategies, potentially increasing compliance costs or gaming behavior; conversely, reliable detectors could deter greenwashing and align incentives toward substantive disclosure.
  • Research agenda: Quantify macroeconomic impacts of erroneous AI-based ESG signals (false positives/negatives) on asset prices, cost of capital, and real investment; evaluate cost-benefit of investing in domain-specific model development vs. manual assurance.

Suggested next steps (technical and policy): - Develop multilingual, ESRS-aligned training datasets and human-labeled visual greenwashing corpora. - Fine-tune or build models with explicit ESRS taxonomies and explainability modules to separate presentation quality from substantive disclosure. - Create standardized evaluation protocols and public benchmarks for ESG greenwashing detection. - Policymakers should avoid relying solely on general-purpose multimodal AI for compliance or enforcement until domain-specific reliability is demonstrated.

Assessment

Paper Typecorrelational Evidence Strengthlow — Small purposive sample (20 Polish ESG reports), absence of independent ground-truth labels for visual greenwashing, single-country/non-English context, potential prompt- and model-version sensitivity, and extremely low inter-tool agreement limit confidence that results generalize or identify true detection performance. Methods Rigorlow — The study uses a standardized prompt and benchmarks against ESRS-aligned national rankings, which is appropriate, but it relies on a purposive small sample, lacks robustness checks (prompt sensitivity, model randomness, alternative labelers), provides no external validation of labels, and does not address temporal/model-update issues—reducing credibility of methodological claims. SamplePurposively selected sample of 20 Polish corporate ESG reports; each report evaluated for visual greenwashing by three multimodal AI systems (ChatGPT 5.1, Claude 4.5 Sonnet, Gemini 2.5 Flash) using a single standardized auditing prompt; firm sustainability performance benchmarked using ESRS-aligned scores from the national "Ranking ESG". Themesgovernance adoption GeneralizabilitySmall purposive sample (n=20) limits statistical power and representativeness, Single-country, non-English (Polish) reporting context may not generalize to other languages or regulatory regimes, Only three proprietary model versions tested; results may change with model updates or different models, Findings specific to visual elements; textual greenwashing and multimodal interactions may behave differently, Prompt formulation and implementation details may drive results (prompt sensitivity), Sector mix of sampled firms may not represent broader corporate population

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The rapid expansion of sustainability reporting under the EU Corporate Sustainability Reporting Directive (CSRD) has intensified concerns about greenwashing, particularly in visual communication within ESG reports. Other negative concerns about greenwashing in ESG visual communication
Reading fidelity high
Study strength low
not reported
0.15
Recent advances in multimodal artificial intelligence offer new possibilities for automated detection of visual greenwashing, yet their reliability in non-English corporate reporting contexts remains unclear. Other mixed reliability of multimodal AI in non-English (Polish) corporate reporting
Reading fidelity high
Study strength low
not reported
0.15
This study evaluates the greenwashing detection capabilities of three leading multimodal AI systems—ChatGPT 5.1, Claude 4.5 Sonnet, and Gemini 2.5 Flash—using a purposively selected sample of 20 Polish ESG reports benchmarked against ESRS-aligned performance scores from the national 'Ranking ESG'. Decision Quality null_result AI greenwashing detection capability compared to ESRS-aligned benchmark
Reading fidelity high
Study strength medium
n=20
0.3
A standardized auditing prompt was applied across all tools to generate comparable assessments of visual greenwashing. Other null_result consistency of assessment procedure across models (use of a standardized prompt)
Reading fidelity high
Study strength medium
n=20
0.3
Contrary to theoretical expectations and all four hypotheses, the models did not demonstrate negative correlations between firm performance and AI-detected greenwashing; instead, high-performing firms frequently received higher greenwashing scores. Decision Quality positive correlation between ESRS-aligned firm performance and AI-detected greenwashing scores
Reading fidelity high
Study strength medium
n=20
0.3
Dimensional analyses showed inconsistent and often contradictory evaluations across Environmental, Social, and Governance pillars. Decision Quality mixed AI greenwashing assessments by ESG pillar (E, S, G)
Reading fidelity high
Study strength medium
n=20
0.3
Inter-tool reliability proved extremely low (Krippendorff’s α ≈ 0). Decision Quality null_result inter-tool agreement (Krippendorff's α) on greenwashing assessments
Reading fidelity high
Study strength medium
n=20
Krippendorff’s α ≈ 0
0.3
These findings indicate that current multimodal AI systems conflate communication sophistication with deceptive intent and lack sufficient contextual understanding for ESG assurance. Decision Quality negative validity of multimodal AI systems for ESG greenwashing detection (confounding by communication sophistication; contextual understanding capability)
Reading fidelity medium
Study strength low
n=20
0.09
The study highlights significant methodological limitations and outlines directions for developing domain-specific, ESRS-aligned AI tools for greenwashing detection. Innovation Output positive recommendation for future tool development (domain-specific, ESRS-aligned AI)
Reading fidelity high
Study strength low
n=20
0.15

Notes