The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Consumer legal AIs confidently get Indian contract law wrong: a 60-case audit finds large high-confidence error rates (Meta AI 31.7%, Perplexity 15.0%, ChatGPT 6.7%) especially on post-2018 Specific Relief issues, while 71% of law students report no formal training in ethical AI use, leaving courts and practitioners exposed to hallucinated citations.

Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?
Angel Mary John, Vipin Kumar Singh, Jerrin Thomas Panachakel · August 21, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Angel Mary John unresolved corpus identity
  2. Vipin Kumar Singh unresolved corpus identity
  3. Jerrin Thomas Panachakel unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Angel Mary John provider ID
  2. Vipin Kumar Singh provider ID
  3. J. T. Panachakel provider ID
A dual audit finds that leading consumer legal AIs often assert incorrect Indian-contract verdicts with very high self-reported confidence (HCER notably 31.7% for Meta AI) while a survey of 380 law students reveals limited formal AI training and reactive rather than proactive verification practices.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized 'precedent overfitting' bias. Phase I of our socio-technical audit tested ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the shift toward statutory enforcement of specific performance. We introduce the High-Confidence Error Rate (HCER) to quantify incorrect verdicts delivered with dangerous certainty (>= 9 on a 1-10 scale). All models struggled with statutory updates. Meta AI proved most vulnerable (31.7% HCER), frequently misapplying pre-amendment rules with a 9.1/10 mean confidence, followed by Perplexity (15.0%) and ChatGPT (6.7%). Phase II investigated human vulnerability to this overconfidence via a survey of Indian law students (N=380). Verification often functions as a reactive adaptation to machine hallucinations: students encountering fabricated citations reported higher verification scores (4.2/5) than those with no such encounters (2.8/5). Furthermore, while 81.6% knew submitting hallucinated cases can lead to contempt-of-court, 71.1% received no formal training on ethical AI use. We propose shifting toward adversarial legal research pedagogy and implementing source-grounded verification architectures to prevent systemic professional negligence.

Summary

Main Finding

The paper uncovers a domain-specific overconfidence in legal LLMs—termed the "inertia of confidence"—where models return incorrect legal conclusions with near-maximal certainty. This behavior appears linked to a hypothesized algorithmic bias called "precedent overfitting," in which models overweight historical pre-amendment case law versus recent statutory changes. In a dual audit (technical + student survey) the authors quantify the phenomenon using a new metric, the High-Confidence Error Rate (HCER): Meta AI showed HCER = 31.7%, Perplexity = 15.0%, and ChatGPT (GPT-5.2) = 6.7% on a specialized 60-case Indian contract law battery; errors concentrate on scenarios testing the Specific Relief (Amendment) Act, 2018.

Key Points

  • Inertia of confidence: LLMs often maintain near-maximum self-reported confidence (scale 1–10) while producing incorrect legal verdicts, especially where modern statutes override voluminous historical precedent.
  • Precedent overfitting: Proposed algorithmic bias where past judicial corpus (pre-amendment decisions) dominates model outputs, causing reversion to outdated law despite training data including newer statutes.
  • HCER (High-Confidence Error Rate): Percentage of model responses that are incorrect and reported with high confidence (threshold used: ≥9/10). HCER results (60-case battery, consumer interfaces):
    • Meta AI: 31.7% HCER; mean confidence ≈ 9.1/10
    • Perplexity (Sonar, RAG): 15.0% HCER
    • ChatGPT (GPT-5.2): 6.7% HCER
  • Performance pattern: Models performed well (80–100% accuracy) on core doctrinal controls (offers, acceptance, capacity), but accuracy dropped markedly for Specific Relief / 2018 Amendment cases (ChatGPT 70%, Perplexity 60%, Meta AI 50%).
  • Human side (survey of N = 380 Indian LLB students):
    • Many students lack formal AI-ethics/training (71.1% reported none).
    • Students with prior encounters of fabricated citations reported higher verification frequency (mean verification score 4.2/5) than students with no such encounters (2.8/5) — evidence of reactive rather than proactive verification.
    • 81.6% aware that submitting hallucinated cases can lead to contempt-of-court, yet institutional preparation is weak.
  • Policy recommendations: adversarial legal research pedagogy, mandatory source-grounded verification architectures for legal AI, and institutional training to avoid systemic professional negligence.

Data & Methods

  • Dual-phase socio-technical audit:
    • Phase I — Algorithmic audit
      • Black-box evaluation via consumer-facing interfaces (to reflect real user exposure):
        • ChatGPT (GPT-5.2) via OpenAI web UI (default parameters)
        • Meta AI via WhatsApp consumer interface
        • Perplexity AI (Sonar) via free-tier web interface (RAG/live retrieval)
      • 60-case "judicial agent" battery focused on Indian Contract Act, 1872 and Specific Relief (Amendment) Act, 2018:
        • Six topical groups (10 cases each): Offer/Acceptance; Capacity/Consent; Consideration/Privity; Discharge/Frustration; Damages/Terms; Specific Relief & 2018 Amendments (trap cases).
        • Factual matrices anonymized to avoid direct retrieval of landmark case names.
      • Standardized prompt: models instructed to act as a senior jurist; required outputs: definitive verdict, supporting statutory authority, self-assessed confidence (1–10).
      • HCER defined as proportion of incorrect verdicts delivered with confidence ≥9 (manual aggregation and evaluation of correctness against ground truth by authors).
      • Temporal reasoning probe: cases constructed to stress retrospective vs prospective application of the 2018 Amendment (e.g., uncertainty after Katta Sujatha Reddy developments).
      • Bias control: uniform prompts, anonymization, black-box consumer interfaces.
    • Phase II — Human survey
      • Cross-sectional primary survey of N = 380 LLB students (purposive convenience sampling across institution types and years).
      • Instrument: 10-question Google Form covering usage, hallucination encounters, institutional training, verification habits, and career anxiety.
      • Key variables: Hallucination Exposure, Verification Frequency (Likert), Institutional Training Status (binary), Liability Awareness.
      • Ethical measures: informed consent, anonymity, no identifying data collected.
  • Limitations noted by authors:
    • Black-box design prevents mechanistic attribution (cannot inspect model weights, retrieval rankings).
    • Prompt framing (judicial persona) may induce response pressure and affect self-reported confidence.
    • Non-probability survey sampling limits generalizability.
    • HCER threshold choice (≥9) is operational and may vary in other studies.

Implications for AI Economics

  • Market for verification and provenance tools: High HCER on legal tasks, especially in statutory-update contexts, creates strong economic demand for verifiable, source-grounded RAG systems, citation provenance services, and third-party legal-AI auditing—supporting investment and competition in "trustworthy legal AI" startups and modules.
  • Liability externalities and cost of error: Overconfident AI errors impose negative externalities (court sanctions, malpractice risk). This raises expected liability costs for legal practitioners and institutions, which will be internalized through higher fees, mandatory verification steps, and professional indemnity insurance premiums. Pricing models for legal services are likely to adjust upward to cover verification overhead.
  • Information asymmetry & reputation risk: AI outputs that appear authoritative but are incorrect exacerbate information asymmetry between providers and consumers. Firms that can credibly certify low HCER (via audits, certifications) will gain competitive advantage; reputational risk from high-profile AI-induced errors will increase the value of certified, audited models.
  • Labor and productivity effects: Short-term productivity gains from AI adoption could be offset by verification burdens and litigation risk (a "productivity–liability tradeoff"). Demand for skilled verification labor (paralegals, AI-auditors) will rise; complementary human capital may become more valuable than raw automation.
  • Regulatory and certification markets: Findings strengthen the case for mandatory auditing standards, HCER-like risk metrics, model disclosure requirements, and certification regimes for legal-AI tools—creating regulatory compliance costs but also new auditing service markets.
  • Education and human capital investment: Weak institutional training implies a market failure in complementarities between AI and users. Investment in adversarial legal-research pedagogy and mandated AI literacy creates positive externalities and reduces systemic risk; universities and law firms will need to invest in curricula and training programs.
  • Insurance and contract design: Insurers will develop contract terms and pricing reflecting HCER exposure and provenance guarantees. Contracts between legal-AI vendors and buyers will likely include warranties, indemnities, and audit clauses—shaping platform economics and bargaining power.
  • Recommendations for economic policy and adoption:
    • Encourage procurement of legal-AI systems with provable provenance and measurable HCER; tie public procurement (e.g., court tools) to audit certification.
    • Subsidize training and verification infrastructure in public-interest legal services to preserve access-to-justice gains while managing error risk.
    • Support creation of market standards (HCER reporting, provenance APIs) to reduce search and information costs and accelerate efficient adoption.

Overall, the paper highlights a measurable, model-level risk (high-confidence legal errors concentrated where statutes change) with direct economic consequences for legal service markets, product-design incentives, regulatory compliance costs, and the structure of labor complementarity in an AI-augmented legal economy.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides quantitative, reproducible-sounding metrics (a 60-case model benchmark with HCER and a 380-student survey) that document meaningful patterns, but the black-box model access, single-task benchmark, limited model set, and non-probability student sampling limit causal claims and external validity. Methods Rigormedium — Strengths include a standardized prompt/persona, an anonymized 60-case battery focused on a concrete statutory shift, and explicit confidence elicitation; weaknesses include black-box access (no access to retrieval/weights), potential prompt-induced bias, limited model coverage, unclear adjudication procedures (e.g., inter-rater coding/reliability), convenience sampling for the survey, and limited robustness checks or statistical inference. SamplePhase I: Black-box evaluation (consumer-facing interfaces) of three LLM systems—ChatGPT (GPT-5.2 via official web UI), Meta AI (WhatsApp interface), and Perplexity AI (Sonar, RAG-based)—on a 60-case anonymized benchmark derived from Indian Contract Act, 1872 and the Specific Relief (Amendment) Act, 2018; models were required to give a verdict, supporting authority, and self-rated confidence (1–10). Data collection occurred in Q1 2026. Phase II: Cross-sectional purposive convenience sample of N = 380 undergraduate LLB students across Indian National Law Universities, private and state-affiliated law schools; self-administered Google Form with 10 structured questions on AI usage, hallucination exposure, verification frequency, training, and attitudes. Themesgovernance human_ai_collab skills_training GeneralizabilityResults limited to three consumer-facing model interfaces and may not extend to other models or API configurations (temperature/retrieval tweaks)., Benchmark focused narrowly on Indian contract law and the 2018 Specific Relief amendment, so findings may not generalize to other legal domains or jurisdictions., Black-box design prevents attributing errors to training data vs. retrieval vs. prompting; thus mechanistic claims (e.g., 'precedent overfitting') remain hypothetical., Student survey used purposive convenience sampling and self-reports, limiting population representativeness and causal inference., Prompt framing (judicial persona and forced definitive verdict) likely increased confidence scores and may not reflect typical user interactions.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Across the 60-case Indian legal battery, accuracy on Specific Relief and the 2018 Amendments cases (Cases 51–60) was 70% for ChatGPT, 60% for Perplexity, and 50% for Meta AI. Decision Quality positive Accuracy of legal verdicts on cases involving the 2018 statutory amendments
Reading fidelity high
Study strength medium
n=60
70% for ChatGPT; 60% for Perplexity; 50% for Meta AI
0.18
ChatGPT achieved 100% accuracy in both the Offer, Acceptance & Communication category and the Capacity, Consent & Formation category. Decision Quality positive Legal verdict accuracy on foundational contract-law cases
Reading fidelity high
Study strength medium
n=20
100% accuracy in both categories
0.18
Meta AI had the highest High-Confidence Error Rate, with 31.7% of its tested verdicts being incorrect while receiving a confidence score of at least 9 out of 10; Perplexity had a rate of 15.0% and ChatGPT 6.7%. Error Rate negative Rate of incorrect legal verdicts delivered with high self-reported confidence
Reading fidelity high
Study strength medium
n=60
Meta AI: 31.7%; Perplexity: 15.0%; ChatGPT: 6.7%
0.18
Meta AI delivered high-confidence errors while reporting a near-perfect mean confidence score of 9.1 out of 10. Decision Quality negative Calibration between legal-verdict correctness and model confidence
Reading fidelity high
Study strength medium
n=60
mean confidence score of 9.1/10
0.18
The observed pattern of lower accuracy on modern statutory updates is consistent with the paper’s hypothesis of ‘precedent overfitting,’ but the study does not establish that precedent overfitting is the causal mechanism. Decision Quality mixed Pattern of legal reasoning errors across historical principles versus modern statutory updates
Reading fidelity high
Study strength low
n=60
0.09
Among surveyed Indian law students, those reporting multiple prior encounters with fabricated legal citations had a higher mean manual-verification score than those reporting no such encounters: 4.2/5 versus 2.8/5. Ai Safety And Ethics positive Frequency of manually verifying AI-generated legal citations
Reading fidelity high
Study strength low
n=380
4.2/5 versus 2.8/5
0.09
A majority of surveyed Indian law students, 71.1%, reported receiving no formal training on the ethical use of AI. Training Effectiveness negative Receipt of formal institutional training on ethical AI use
Reading fidelity high
Study strength low
n=380
71.1% reported receiving no formal training
0.09
Among surveyed Indian law students, 81.6% reported awareness that submitting hallucinated cases to an Indian court can lead to contempt-of-court consequences. Governance And Regulation positive Awareness of judicial liability for submitting hallucinated cases
Reading fidelity high
Study strength low
n=380
81.6% reported awareness
0.09
The surveyed law students reported a mean job-displacement anxiety score of 3.34 out of 5. Job Displacement negative Self-reported anxiety about AI-related job displacement
Reading fidelity high
Study strength low
n=380
µ = 3.34/5
0.09
The study recommends adversarial legal-research pedagogy and mandatory source-grounded verification architectures for legal AI systems to reduce reliance on confident but hallucinated outputs. Governance And Regulation positive Proposed mitigation of legal-AI hallucination and professional negligence risks
Reading fidelity high
Study strength speculative
n=380
0.03

Notes