The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Publicly reported benchmarks show Indian models score well on older saturated tests but underparticipate in newer agentic and domain-specific evaluations; the paper argues many apparent capability shortfalls may reflect gaps in benchmarking and disclosure rather than true technical weakness and proposes a Benchmark Maturity Index to guide national monitoring.

Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework
Avinash Agarwal, Vridhi Jain · August 12, 2026
arxiv descriptive low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Avinash Agarwal unresolved corpus identity
  2. Vridhi Jain unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Avinash Agarwal provider ID
  2. Vridhi Jain provider ID
Using publicly reported benchmark results across eight capability domains, the paper finds Indian foundation models show strong performance on saturated benchmarks but participate far less in newer, agentic and domain-specialized evaluations, and proposes a four-dimension Benchmark Maturity Index to distinguish capability gaps from evaluation-ecosystem gaps.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Assessing the progress of such national ecosystems is complicated by inconsistent benchmark reporting, proprietary evaluation methodologies, and rapidly evolving model releases. This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models, across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI and computer use, cybersecurity, vision and image understanding, video and multimodal understanding, scientific research, and Indic language capability. Using only publicly reported benchmark results, we find that Indian models achieve strong scores on established benchmarks such as MMLU and MATH-500. However, these benchmarks are now widely regarded as saturated, and frontier developers no longer report them. Indian models participate far less frequently in newer, agentic, and domain-specialized evaluations. Benchmark participation is also highly uneven across Indian organizations. Among the models surveyed, Sarvam AI reports the broadest benchmark coverage by a substantial margin. We propose an exploratory four-dimension Benchmark Maturity Index (BMI), scoring each capability domain on standardization, participation, independent verification, and national coverage. We show that the BMI refines, and in some cases revises, the maturity judgments that a purely descriptive review would produce. We argue that many apparent capability gaps in the public record cannot be distinguished, on available evidence, from evaluation-ecosystem gaps. This has direct implications for how national AI programs should design monitoring and funding criteria.

Summary

Main Finding

Publicly reported benchmark results show Indian foundation models demonstrate strong performance on older, now-saturated benchmarks (e.g., MMLU, MATH-500, HumanEval/MBPP). However, Indian models participate far less in newer, agentic, and domain-specialized evaluations that better discriminate frontier capability. Benchmark participation and disclosure are uneven across Indian organizations (Sarvam AI has the broadest public coverage). The paper introduces a Benchmark Maturity Index (BMI) to distinguish true capability gaps from gaps in the public evaluation ecosystem — a distinction that matters for monitoring, funding, and national AI policy.

Key Points

  • Benchmark snapshot (August 2026): Indian models score highly on several established benchmarks but those benchmarks are widely regarded as saturated at the frontier.
  • Frontier developers increasingly stop reporting on saturated benchmarks, shifting attention to harder, agentic, or domain-specific evaluations; Indian models rarely appear on these newer benchmarks.
  • Benchmark disclosure is highly uneven among IndiaAI-supported organizations: only Sarvam AI provides broad multi-domain public results; many supported groups publish little or no standardized benchmark evidence.
  • The paper proposes a three-tier working definition of "Indian-developed" models:
    • Tier 1: fully indigenous (trained from scratch on India-controlled compute and significant India-sourced data).
    • Tier 2: India-led with global components.
    • Tier 3: India-adapted foreign base models (fine-tuned/adapted).
  • Introduces a four-dimension Benchmark Maturity Index (BMI) at the capability-domain level, scoring:
    • Standardization (existence of accepted benchmarks),
    • Participation (number/representativeness of models evaluated),
    • Independent verification (third-party/leaderboard confirmation),
    • National coverage (domestic developer participation and disclosure).
  • The BMI refines and sometimes revises maturity judgments from qualitative review alone; many apparent capability gaps are equally plausibly explained by evaluation-ecosystem gaps.
  • Key caveat: analysis uses only publicly reported scores (developer-reported vs independently verified are distinguished) and is therefore a snapshot of the public record, not the full state of capability.

Data & Methods

  • Model groups compared:
    • Frontier: GPT-5.6, Claude Opus 5, Kimi K3, Qwen3.8-Max (top models on third-party leaderboards with public technical reports).
    • Comparable-scale: models defined by active-parameter range (~12–50B active): Inkling (MoE), Qwen3.6-27B, Nemotron 3 Super.
    • Indian models: Sarvam-105B, Sarvam-30B, Param2 (BharatGen), Krutrim-2; plus domain-specific Indian systems (e.g., Shodh AI’s Project Skanda, Avataar AI’s Varya).
  • Capability domains (eight): general-purpose reasoning/knowledge, coding/software engineering, agentic AI/computer use, cybersecurity, vision/image understanding, video/multimodal understanding, scientific research, Indic language capability.
  • Benchmarks used (examples): MMLU, GPQA Diamond, HLE, MATH-500, ARC-AGI; HumanEval, MBPP, LiveCodeBench, SWE-bench Pro, Terminal-Bench 2.1; domain-specific leaderboards where applicable.
  • Data sources: only publicly available benchmark results from technical reports, model cards, and third‑party leaderboards. No independent evaluations or re-runs were performed by the authors.
  • Distinctions made: developer-reported vs independently verified scores; BMI participation dimension scored against the four-model frontier sample (sensitivity noted).
  • Timebound: data collection and analysis reflect the public record as of August 2026.

Implications for AI Economics

  • Measurement and inference

    • Public benchmark scores are an imperfect proxy for national AI capability. Relying on public benchmarks risks under- or mis-estimating capability where disclosure is uneven.
    • Benchmark-saturation effects mean high scores on older tests do not imply frontier competitiveness; economic assessments should weight benchmark discriminativity and recency.
    • Evaluation-ecosystem maturity (availability, standardization, independent verification) must be modeled separately from model capability in national-level productivity and capability metrics.
  • Policy and funding design

    • Monitoring and funding criteria for publicly supported model programs should require standardized benchmark disclosure and independent verification to avoid conflating reporting gaps with capability gaps.
    • Differentiate funding by development tier: Tier 1 efforts may need sustained public investment in compute, data, and engineering capacity; Tier 2/3 work may require targeted incentives (e.g., fine-tuning/data curation).
    • Support participation in harder, agentic, and domain-specific benchmarks (which better predict real-world economic impact) via grants, benchmarking credits, or centralized evaluation infrastructure.
  • Strategic and market effects

    • Transparent public benchmarking is a signal to purchasers, partners, and labor markets; uneven disclosure may distort procurement decisions and private investment.
    • National benchmarking infrastructure and standardized public leaderboards can increase the visibility and comparability of domestic models, reducing information asymmetries that affect market entry, procurement, and international collaboration.
    • Investing in independent benchmarking capacity (national evaluation hubs, standardized suites for Indic languages and domain tasks) creates public goods that lower assessment costs for both public and private actors and improves the accuracy of economic indicators tied to AI capability.
  • Research and workforce development

    • Emphasize benchmarks that reflect multi-step, agentic, and domain-specific tasks to better align evaluation with economic use-cases (software engineering, scientific discovery, multimodal services).
    • Public benchmarking programs can stimulate skills development, standardize evaluation practice, and create datasets and tooling with spillovers across academia and industry.

Recommendations (concise) - Require public disclosure of standardized benchmark suites and encourage independent verification for recipients of public funds. - Invest in national benchmarking infrastructure and updated benchmark suites that include agentic and domain-specialized evaluations, plus strong Indic-language tests. - Use the BMI (or a refined version) as part of funding/monitoring dashboards to distinguish capability from evaluation-maturity gaps. - Tailor public funding by model tier: prioritize compute/data for Tier 1 and incentivize participation in advanced benchmarks for all tiers.

Limitations to bear in mind - The study uses only public reports (snapshot August 2026); non-disclosure does not imply lack of capability. - Active parameter counts and architecture heterogeneity limit strict comparability across models. - BMI is exploratory and intended as a framework for further development rather than a validated index.

Assessment

Paper Typedescriptive Evidence Strengthlow — The paper compiles and compares only publicly reported benchmark scores without running independent evaluations; results are subject to self-reporting bias, selective disclosure, benchmark saturation, and evolving leaderboards, so the evidence cannot reliably establish true capability gaps or frontier comparability. Methods Rigormedium — The authors use a transparent, documented selection procedure (three model groups, eight capability domains), distinguish developer-reported vs independently verified scores, and introduce a structured Benchmark Maturity Index; however, they do not reproduce benchmarks, rely on heterogeneous public disclosures, and the BMI is exploratory and not validated, limiting methodological strength. SampleSnapshot (August 2026) of publicly reported benchmark scores extracted from model technical reports, model cards, and third-party leaderboards for three model groups: (1) four frontier models (GPT-5.6, Claude Opus 5, Kimi K3, Qwen3.8-Max), (2) three comparable-scale global models defined by active-parameter range (Inkling, Qwen3.6-27B, Nemotron 3 Super), and (3) representative Indian models (Sarvam-105B, Sarvam-30B, Param2/BharatGen, Krutrim-2, plus domain-specific Indian systems); no original evaluations were run and only publicly disclosed metrics were used. Themesgovernance innovation GeneralizabilityConditional on public disclosure: models or results not publicly reported are invisible and may bias conclusions, Snapshot in time (August 2026); benchmark scores and model releases change rapidly, Heterogeneous benchmark selection and harness configurations across developers limit direct comparability, Active-parameter and architecture heterogeneity (MoE vs dense) complicates scale-based comparisons, Benchmark saturation: high scores on older benchmarks (MMLU, MATH-500, HumanEval) are less informative about frontier performance, Findings do not generalize to deployment outcomes (cost, latency), safety/alignment, or real-world economic impacts

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The publicly benchmarked Indian models achieve strong scores on MMLU and MATH-500, but these benchmarks are saturated and do not establish frontier competitiveness. Output Quality mixed Benchmark performance on general-purpose knowledge and mathematical reasoning
Reading fidelity high
Study strength medium
n=11
0.18
On GPQA Diamond, Sarvam-105B scores 78.7, trailing the surveyed frontier leader by roughly 15 points and the best comparable-scale model by 8.5 points. Output Quality negative General-purpose reasoning performance on GPQA Diamond
Reading fidelity high
Study strength medium
n=11
roughly 15 points behind the frontier leader; 8.5 points behind the comparable-scale leader
0.18
On HLE with tools, the best comparable-scale model substantially outperforms Sarvam-105B, scoring 46.0 versus 11.2. Output Quality negative Hard reasoning performance on Humanity's Last Exam with tools
Reading fidelity high
Study strength medium
n=11
34.8-point score difference
0.18
No Indian model reports a score on ARC-AGI-2 or ARC-AGI-3. Adoption Rate null_result Public benchmark participation in advanced reasoning evaluations
Reading fidelity high
Study strength low
n=4
0.09
Indian models report strong scores on foundational coding benchmarks, including HumanEval, MBPP, and LiveCodeBench. Developer Productivity positive Foundational code-generation benchmark performance
Reading fidelity high
Study strength medium
n=4
92.1 on HumanEval; 92.7 on MBPP; 71.7 on LiveCodeBench
0.18
No Indian model reports a result on any of the four agentic or multi-step software-engineering benchmarks reviewed. Adoption Rate null_result Public participation in agentic and multi-step software-engineering evaluations
Reading fidelity high
Study strength low
n=4
0.09
The available public evidence shows Indian benchmark participation concentrated in foundational coding, while it remains absent from the reviewed agentic coding benchmarks. Adoption Rate mixed Distribution of public benchmark participation across coding-evaluation types
Reading fidelity high
Study strength low
n=4
0.09
The paper cannot determine from available evidence whether the absence of Indian results on agentic coding benchmarks reflects a genuine capability gap or only a reporting gap. Ai Safety And Ethics mixed Interpretability of public benchmark nonparticipation as capability versus disclosure
Reading fidelity high
Study strength low
n=4
0.09
Among the 12 organizations supported by the IndiaAI Innovation Centre's Foundation Models pillar, only Sarvam AI publishes benchmark results across multiple capability domains; BharatGen/Param2 reports limited results, and the remaining 10 have not disclosed standardized scores within the assessed domains. Adoption Rate negative Public benchmark disclosure and evaluation coverage among supported organizations
Reading fidelity high
Study strength medium
n=12
1 of 12 organizations with broad multi-domain benchmark coverage; 10 of 12 with no disclosed standardized scores in assessed domains
0.18
Sarvam AI reports the broadest benchmark coverage among the surveyed Indian organizations by a substantial margin. Adoption Rate positive Breadth of public benchmark participation across capability domains
Reading fidelity high
Study strength medium
n=12
0.18
The proposed Benchmark Maturity Index refines, and in two cases revises, maturity judgments produced by a purely qualitative review. Governance And Regulation positive Maturity assessment of the national AI evaluation ecosystem
Reading fidelity high
Study strength low
n=8
two cases revised
0.09

Notes