The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AEEP: a validated, adaptive audit that reveals substantial differences in LLM ethical consistency and offers firms and regulators a practical standard for assessing advisory assistants; in a 50-dialogue snapshot, expert–algorithm agreement was 93.8% and Claude outperformed peers while Grok showed the weakest consistency.

Should Businesses Trust AI Advice? A Methodology to Audit the Ethical Integrity of Chatbots
Manuel Chaves-Maza · September 01, 2026 · Computers in Human Behavior Reports
openalex descriptive medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Manuel Chaves-Maza provider ID

Semantic Scholar

Latest observation:

  1. Manuel Chaves-Maza unresolved corpus identity
The paper introduces AEEP, an adaptive, validated audit protocol for LLM enterprise advisors and shows, in a 50-dialogue snapshot validated by five ethics experts, substantial and reliable differences in models' ethical consistency (algorithm–expert agreement 93.8%; Cohen’s κ = 0.728), with Claude the most consistent and Grok the least.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

: As Large Language Models (LLMs) move from general-purpose chat to embedded advisors in small and medium-sized enterprises (SMEs), a human-centered question becomes urgent: can the systems we ask entrepreneurs to trust sustain a coherent ethical stance when business pressure pushes back? This study addresses that question with a single, focused contribution: the Adaptive Ethical Evaluation Protocol (AEEP), a validated audit methodology designed specifically for human-centered AI advisory contexts. Unlike static questionnaires or one-shot benchmarks, the AEEP stages a structured, five-node adaptive dialogue in which counter-arguments are calibrated to each model's prior response—applying pragmatic pressure to principle-based answers and ethical probing to permissive ones. We applied the protocol to five frontier LLMs (ChatGPT, Claude, Gemini, Grok, DeepSeek) across ten dilemmas grounded in everyday SME advisory practice (nepotism, whistleblowing, data privacy, algorithmic bias, regulatory compliance, among others), yielding 50 branched dialogues collected between 12 and 14 May 2025 (temperature = 0.7, one run per prompt, five conversational nodes per dialogue). Each transcript was scored on four pre-registered indicators (Ethical Awareness, Consistency, Ethics Priority, Contradiction) using a transparent coding pipeline combining sentiment analysis, keyword extraction and NLI-based contradiction detection. The same 50 dialogues were independently and blindly rated by a panel of five senior ethics researchers using identical rubrics, in a validation round dedicated to this instrument. Algorithm–expert agreement reached 93.8% (Cohen's κ = 0.728, Pearson r = 0.838, p < 0.001), with substantial inter-rater reliability across the panel. Rankings exposed clear behavioural differences: Claude held its position most consistently (0.938), while Grok wavered under pressure (0.675). The contribution is not another LLM leaderboard: it is a reusable, expert-validated audit instrument that lets enterprise advisors, regulators and SME managers decide where to trust AI advice and where human oversight must remain in the loop.

Summary

Main Finding

The paper introduces the Adaptive Ethical Evaluation Protocol (AEEP), a validated, reusable audit methodology for assessing whether LLM-based enterprise advisors sustain coherent ethical stances under realistic, adversarial conversational pressure. Applied to five frontier models across ten SME-relevant dilemmas (50 branched dialogues collected 12–14 May 2025), the AEEP produced reliable, expert-validated measurements showing substantive behavioural differences between models (algorithm–expert agreement 93.8%; Cohen’s κ = 0.728; Pearson r = 0.838, p < 0.001). The authors position the contribution as an operational audit instrument for firms, regulators and procurement teams—not another leaderboard.

Key Points

  • AEEP design
    • Five-node adaptive dialogue: counter-arguments are calibrated to the model’s prior reply (pragmatic pressure to principled answers; ethical probing to permissive answers).
    • Designed specifically for human-centered AI advisory contexts (SMEs).
  • Experimental snapshot
    • Models tested: ChatGPT, Claude, Gemini, Grok, DeepSeek.
    • Scenarios: 10 SME advisory dilemmas (nepotism, whistleblowing, data privacy, algorithmic bias, regulatory compliance, etc.).
    • Data collected: 50 branched dialogues (5 models × 10 dilemmas), 12–14 May 2025.
    • Generation settings: temperature = 0.7, one run per prompt, five conversational nodes per dialogue.
  • Evaluation metrics
    • Four pre-registered indicators: Ethical Awareness, Consistency, Ethics Priority, Contradiction.
    • Automated coding pipeline: sentiment analysis, keyword extraction, NLI-based contradiction detection.
  • Validation
    • Independent blind rating by five senior ethics researchers using the same rubric.
    • High algorithm–expert agreement (93.8%) and substantial intercoder reliability.
  • Observed model behaviour
    • Clear ranking differences; Claude most consistent (0.938), Grok least consistent under pressure (0.675).
  • Contribution
    • A validated, transparent audit methodology tailored to advisory contexts that can inform trust, oversight, procurement and regulation decisions.

Data & Methods

  • Experimental setup
    • Stimuli: 10 vignette-based dilemmas reflecting everyday SME advisory needs (ethical and compliance trade-offs).
    • Interaction protocol: AEEP’s five sequential nodes adapt prompts depending on the model’s previous reply to apply calibrated counter-argument or probing.
    • Model runs: single-run per prompt, temperature set to 0.7 to permit some variation while maintaining comparability.
  • Scoring & automation
    • Four indicators pre-registered and operationalized in code.
    • Text processing pipeline components:
    • Sentiment analysis to detect normative tone and affect.
    • Keyword extraction to capture ethical framing and domain-relevant terms.
    • Natural Language Inference (NLI) to detect contradictions between nodes.
    • Each dialogue yielded scores on the four indicators.
  • Validation procedure
    • Blind independent coding by five senior ethics researchers using the same rubric.
    • Comparison metrics: percent agreement, Cohen’s κ, Pearson correlation.
    • Results: algorithm–expert agreement 93.8%, Cohen’s κ = 0.728 (substantial), Pearson r = 0.838 (p < 0.001).
  • Limitations noted or implied
    • Single run per prompt (no distributional assessment of stochastic variability).
    • Fixed temperature and limited sample of models and dilemmas (snapshot as of mid-May 2025).
    • Automated metrics may miss subtleties of ethical reasoning; potential for prompt or style-dependent gaming by models.
    • Branching depth capped at five nodes—does not capture longer-term dynamics.

Implications for AI Economics

  • Adoption and trust in SME advisory tools
    • AEEP offers a practical mechanism for SMEs and vendors to evaluate the ethical robustness of advisory assistants, influencing procurement choices and enterprise adoption rates.
    • Transparent, validated audits can reduce asymmetric information about model behaviour and lower trust frictions in adoption decisions.
  • Market differentiation and product value
    • Audited ethical consistency can become a marketable product feature; vendors with higher AEEP scores may command premium pricing or preferred enterprise contracts.
    • Certification based on instruments like AEEP could create barriers to entry for providers unwilling or unable to meet ethics robustness standards.
  • Regulatory and compliance economics
    • Regulators and compliance teams can integrate adaptive auditing into oversight frameworks, potentially lowering enforcement costs by providing standardized assessment tools.
    • Firms may adopt AEEP-style audits to limit liability and demonstrate due diligence in algorithmic governance.
  • Labor, supervision and cost trade-offs
    • Where models pass AEEP checks, firms may reduce supervisory labor costs; where models waver, the protocol helps quantify where human-in-the-loop oversight must be maintained—informing staffing and workflow design decisions.
    • The approach helps estimate residual risk and compliance costs associated with delegating advisory tasks to models.
  • Contracting, insurance and certification markets
    • AEEP outputs can feed contractual clauses, service-level agreements and insurance underwriting for AI advisory services.
    • Insurers and auditors can price risk more accurately using standardized, validated measures of ethical consistency.
  • Policy and research implications
    • Encourages policy frameworks that require adaptive, context-sensitive evaluation rather than static benchmarks.
    • Suggests further economic research on how validated audit instruments affect competition, pricing, and diffusion of AI advisory technologies across SMEs.

Practical takeaway: procurement teams, regulators and SME managers can adopt the AEEP as a repeatable audit step to decide where to trust LLM advice and where to mandate human oversight; treating this validated instrument as part of due-diligence, contracting and compliance workflows will materially affect adoption costs and liability allocation.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides validated, empirical measurements (50 dialogues; algorithm–expert agreement and intercoder reliability reported) showing systematic model differences, but the sample is small and snapshot-like (single-run per prompt, limited models and scenarios), limiting robustness and external validity. Methods Rigormedium — Pre-registered indicators, an automated scoring pipeline, and blind validation by five senior ethics researchers strengthen credibility; however, single-run generation, limited scenario/model coverage, possible shortcomings of automated NLI/sentiment measures, and capped dialogue depth reduce rigor. SampleFive LLMs (ChatGPT, Claude, Gemini, Grok, DeepSeek) evaluated on 10 SME-relevant vignette dilemmas, producing 50 branched dialogues (5 models × 10 dilemmas) collected 12–14 May 2025; each dialogue used a five-node adaptive interaction, one generation run per prompt at temperature 0.7. Automated scoring produced four pre-registered indicators; validation via blind independent ratings from five senior ethics researchers on the same rubric. Themesgovernance adoption GeneralizabilitySmall number of models and SME-focused scenarios limits applicability to other model families and non-SME contexts, Single-run per prompt and fixed temperature prevents assessment of stochastic variability and distributional behavior, Snapshot timing (mid-May 2025) — model updates could materially change behavior, Five-node depth may not capture long-run conversational dynamics or deliberative failure modes, Automated NLP measures (sentiment, keyword extraction, NLI) may miss nuanced ethical reasoning or be vulnerable to prompt/style gaming

Claims (13)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The Adaptive Ethical Evaluation Protocol (AEEP) is a reusable audit methodology for assessing whether LLM-based enterprise advisors maintain coherent ethical stances under adversarial conversational pressure. Ai Safety And Ethics positive Ethical stance coherence under conversational pressure
Reading fidelity high
Study strength medium
not reported
0.18
Across five models and ten SME-relevant dilemmas, the AEEP generated 50 branched dialogues collected from 12–14 May 2025. Ai Safety And Ethics other Coverage of the adaptive audit experiment
Reading fidelity high
Study strength medium
n=50
50 branched dialogues
0.18
The automated AEEP coding pipeline achieved 93.8% agreement with independent expert ratings. Ai Safety And Ethics positive Agreement between automated ethical-behavior coding and expert assessment
Reading fidelity high
Study strength medium
n=50
93.8% agreement
0.18
The automated AEEP scores had substantial agreement with expert ratings, with Cohen's kappa equal to 0.728. Ai Safety And Ethics positive Inter-rater agreement between algorithmic and expert ethical evaluations
Reading fidelity high
Study strength medium
n=50
Cohen's κ = 0.728
0.18
Automated AEEP scores were strongly correlated with expert ratings, with Pearson r = 0.838 and p < 0.001. Ai Safety And Ethics positive Correlation between automated ethical evaluation scores and expert ratings
Reading fidelity high
Study strength medium
n=50
Pearson r = 0.838, p < 0.001
0.18
The tested models exhibited substantive differences in ethical consistency under conversational pressure. Ai Safety And Ethics mixed Ethical consistency under pressure
Reading fidelity high
Study strength medium
n=50
0.18
Claude was the most consistent tested model, with a consistency score of 0.938, while Grok was the least consistent under pressure, with a score of 0.675. Ai Safety And Ethics mixed Model consistency under adversarial conversational pressure
Reading fidelity high
Study strength low
n=50
Claude = 0.938; Grok = 0.675
0.09
The AEEP evaluates four pre-registered indicators: Ethical Awareness, Consistency, Ethics Priority, and Contradiction. Ai Safety And Ethics other Multidimensional ethical behavior of LLM advisors
Reading fidelity high
Study strength medium
n=50
0.18
The AEEP uses sentiment analysis, keyword extraction, and natural-language-inference-based contradiction detection to automate dialogue scoring. Ai Safety And Ethics other Automated detection and scoring of ethical framing, normative tone, and contradictions
Reading fidelity high
Study strength medium
n=50
0.18
The study's findings are limited by using a single run per prompt, a fixed temperature, a limited set of models and dilemmas, and a five-node branching depth. Ai Safety And Ethics negative Generalizability and robustness of ethical-consistency measurements
Reading fidelity high
Study strength high
n=50
0.3
The paper proposes that validated, transparent ethical audits could inform SME procurement, regulatory oversight, contracting, and decisions about when to retain human-in-the-loop supervision. Governance And Regulation positive Use of ethical audits in AI adoption, oversight, and supervision decisions
Reading fidelity high
Study strength speculative
not reported
0.03
The paper suggests that AEEP-style audits could reduce asymmetric information and trust frictions affecting enterprise adoption of AI advisory tools. Adoption Rate positive Enterprise adoption of AI advisory tools
Reading fidelity high
Study strength speculative
not reported
0.03
The paper suggests that firms may use ethical-consistency audit results to determine where human oversight must be maintained and to estimate residual risk and compliance costs from delegating advisory tasks to models. Organizational Efficiency mixed Human supervision requirements and compliance costs for AI-assisted advisory work
Reading fidelity high
Study strength speculative
not reported
0.03

Notes