3 cumulative citations
View corpus contextLeading multimodal AI systems struggle to detect greenwashing in Polish ESG reports, frequently flagging high-performing firms as more deceptive and showing virtually no agreement across models; current off‑the‑shelf tools conflate polished sustainability communication with fraudulent intent and lack the contextual grounding for reliable ESG assurance.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
2 cumulative citations
View corpus contextThe rapid expansion of sustainability reporting under the EU Corporate Sustainability Reporting Directive (CSRD) has intensified concerns about greenwashing, particularly in visual communication within ESG reports. Recent advances in multimodal artificial intelligence offer new possibilities for automated detection, yet their reliability in non-English corporate reporting contexts remains unclear. This study evaluates the greenwashing detection capabilities of three leading multimodal AI systems—ChatGPT 5.1, Claude 4.5 Sonnet, and Gemini 2.5 Flash—using a purposively selected sample of 20 Polish ESG reports benchmarked against ESRS-aligned performance scores from the national “Ranking ESG”. A standardized auditing prompt was applied across all tools to generate comparable assessments of visual greenwashing. Contrary to theoretical expectations and all four hypotheses, the models did not demonstrate negative correlations between performance and AI-detected greenwashing; instead, high-performing firms frequently received higher greenwashing scores. Dimensional analyses showed inconsistent and often contradictory evaluations across Environmental, Social, and Governance pillars, while inter-tool reliability proved extremely low (Krippendorff’s α ≈ 0). These findings indicate that current multimodal AI systems conflate communication sophistication with deceptive intent and lack sufficient contextual understanding for ESG assurance. The study highlights significant methodological limitations and outlines directions for developing domain-specific, ESRS-aligned AI tools for greenwashing detection.
Summary
Main Finding
Multimodal LLMs (ChatGPT 5.1, Claude 4.5 Sonnet, Gemini 2.5 Flash) failed to reliably detect visual greenwashing in non-English (Polish) ESG reports. Rather than showing the expected negative correlation between substantive ESG performance and AI-flagged greenwashing, the tools often flagged higher-performing firms more heavily. Across Environmental, Social, and Governance dimensions the models produced inconsistent and contradictory judgments, and inter-model agreement was effectively nil (Krippendorff’s α ≈ 0). Overall, current general-purpose multimodal systems conflate polished communication with deceptive intent and lack the contextual grounding required for trustworthy ESG assurance.
Key Points
- Models evaluated: ChatGPT 5.1, Claude 4.5 Sonnet, Gemini 2.5 Flash.
- Sample: purposively selected n = 20 Polish ESG reports, benchmarked to ESRS-aligned performance scores from the national “Ranking ESG”.
- Procedure: a standardized auditing prompt applied identically across tools to assess visual greenwashing in reports.
- Main unexpected result: AI greenwashing scores were not negatively correlated with ESRS-aligned performance; high-performing firms frequently received higher greenwashing scores.
- Pillar-level analysis (E, S, G) produced inconsistent and often contradictory results across tools.
- Inter-tool reliability was effectively zero (Krippendorff’s α ≈ 0), indicating no reproducible agreement among systems.
- Interpretation: models appear to equate sophisticated visual/communicative quality with possible deception rather than assessing substantive performance or ESRS-conformant disclosures.
- Study cautions against using off-the-shelf multimodal LLMs for ESG assurance in non-English contexts.
Data & Methods
- Data: 20 Polish corporate ESG reports selected purposively to span a range of ESRS-aligned performance scores from the national “Ranking ESG”.
- Benchmark: ESRS-aligned performance scores from Ranking ESG served as the substantive-performance benchmark.
- Tools: three leading multimodal AI systems—ChatGPT 5.1, Claude 4.5 Sonnet, Gemini 2.5 Flash—accessed contemporaneously and prompted with the same standardized audit instruction.
- Prompting: a single, standardized visual-audit prompt was used to generate comparable greenwashing assessments across tools; outputs were mapped to a common scoring rubric.
- Analyses:
- Correlation tests between AI greenwashing scores and ESRS-aligned performance scores (hypothesized negative correlation).
- Dimensional breakdowns for Environmental, Social, and Governance pillars.
- Reliability analysis across models using Krippendorff’s α.
- Key methodological limitations noted by authors:
- Small, purposive, single-country (Poland) sample limits generalizability.
- Non-English reporting poses known challenges for models trained predominantly on English data.
- A single standardized prompt may not capture optimal prompting strategies per model.
- Off-the-shelf models not fine-tuned on ESRS taxonomy or domain-specific labeled datasets.
Implications for AI Economics
- Market signals and capital allocation: If general-purpose multimodal AIs flag high-performing firms as greenwashing (false positives), automated screening tools could distort investor signals and lead to mispricing or misallocation of capital.
- Demand for domain-specific assurance tools: Findings increase economic value for specialized, ESRS-aligned multimodal models or classifiers trained on multilingual, audit-labeled corpora—creating an opportunity for vendors and service providers in AI assurance markets.
- Regulatory and compliance risk: Regulators and standard setters should be cautious about relying on off-the-shelf AI for ESG verification; poor reliability could produce regulatory errors and legal risks for firms and auditors.
- Productivity and cost of assurance: While automated detection promises cost savings, current tools may raise verification costs (manual review of false positives/negatives) until domain-tuned systems reach acceptable accuracy and reliability.
- Need for standardized benchmarks and public datasets: Economically efficient development of trustworthy tools requires labeled, multilingual, ESRS-aligned datasets and open benchmarks to reduce information asymmetries between AI vendors and users.
- Incentives and strategic behavior: If firms learn that polished presentation triggers automated scrutiny, they may alter communication strategies, potentially increasing compliance costs or gaming behavior; conversely, reliable detectors could deter greenwashing and align incentives toward substantive disclosure.
- Research agenda: Quantify macroeconomic impacts of erroneous AI-based ESG signals (false positives/negatives) on asset prices, cost of capital, and real investment; evaluate cost-benefit of investing in domain-specific model development vs. manual assurance.
Suggested next steps (technical and policy): - Develop multilingual, ESRS-aligned training datasets and human-labeled visual greenwashing corpora. - Fine-tune or build models with explicit ESRS taxonomies and explainability modules to separate presentation quality from substantive disclosure. - Create standardized evaluation protocols and public benchmarks for ESG greenwashing detection. - Policymakers should avoid relying solely on general-purpose multimodal AI for compliance or enforcement until domain-specific reliability is demonstrated.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The rapid expansion of sustainability reporting under the EU Corporate Sustainability Reporting Directive (CSRD) has intensified concerns about greenwashing, particularly in visual communication within ESG reports. Other | negative | concerns about greenwashing in ESG visual communication |
Reading fidelity
high
Study strength
low
|
not reported
|
| Recent advances in multimodal artificial intelligence offer new possibilities for automated detection of visual greenwashing, yet their reliability in non-English corporate reporting contexts remains unclear. Other | mixed | reliability of multimodal AI in non-English (Polish) corporate reporting |
Reading fidelity
high
Study strength
low
|
not reported
|
| This study evaluates the greenwashing detection capabilities of three leading multimodal AI systems—ChatGPT 5.1, Claude 4.5 Sonnet, and Gemini 2.5 Flash—using a purposively selected sample of 20 Polish ESG reports benchmarked against ESRS-aligned performance scores from the national 'Ranking ESG'. Decision Quality | null_result | AI greenwashing detection capability compared to ESRS-aligned benchmark |
Reading fidelity
high
Study strength
medium
|
n=20
|
| A standardized auditing prompt was applied across all tools to generate comparable assessments of visual greenwashing. Other | null_result | consistency of assessment procedure across models (use of a standardized prompt) |
Reading fidelity
high
Study strength
medium
|
n=20
|
| Contrary to theoretical expectations and all four hypotheses, the models did not demonstrate negative correlations between firm performance and AI-detected greenwashing; instead, high-performing firms frequently received higher greenwashing scores. Decision Quality | positive | correlation between ESRS-aligned firm performance and AI-detected greenwashing scores |
Reading fidelity
high
Study strength
medium
|
n=20
|
| Dimensional analyses showed inconsistent and often contradictory evaluations across Environmental, Social, and Governance pillars. Decision Quality | mixed | AI greenwashing assessments by ESG pillar (E, S, G) |
Reading fidelity
high
Study strength
medium
|
n=20
|
| Inter-tool reliability proved extremely low (Krippendorff’s α ≈ 0). Decision Quality | null_result | inter-tool agreement (Krippendorff's α) on greenwashing assessments |
Reading fidelity
high
Study strength
medium
|
n=20
Krippendorff’s α ≈ 0
|
| These findings indicate that current multimodal AI systems conflate communication sophistication with deceptive intent and lack sufficient contextual understanding for ESG assurance. Decision Quality | negative | validity of multimodal AI systems for ESG greenwashing detection (confounding by communication sophistication; contextual understanding capability) |
Reading fidelity
medium
Study strength
low
|
n=20
|
| The study highlights significant methodological limitations and outlines directions for developing domain-specific, ESRS-aligned AI tools for greenwashing detection. Innovation Output | positive | recommendation for future tool development (domain-specific, ESRS-aligned AI) |
Reading fidelity
high
Study strength
low
|
n=20
|