0 cumulative citations
View corpus contextPublicly reported benchmarks show Indian models score well on older saturated tests but underparticipate in newer agentic and domain-specific evaluations; the paper argues many apparent capability shortfalls may reflect gaps in benchmarking and disclosure rather than true technical weakness and proposes a Benchmark Maturity Index to guide national monitoring.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual computing. Assessing the progress of such national ecosystems is complicated by inconsistent benchmark reporting, proprietary evaluation methodologies, and rapidly evolving model releases. This paper presents a structured, benchmark-based comparative assessment of publicly benchmarked Indian foundation models against global frontier and comparable-scale models, across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI and computer use, cybersecurity, vision and image understanding, video and multimodal understanding, scientific research, and Indic language capability. Using only publicly reported benchmark results, we find that Indian models achieve strong scores on established benchmarks such as MMLU and MATH-500. However, these benchmarks are now widely regarded as saturated, and frontier developers no longer report them. Indian models participate far less frequently in newer, agentic, and domain-specialized evaluations. Benchmark participation is also highly uneven across Indian organizations. Among the models surveyed, Sarvam AI reports the broadest benchmark coverage by a substantial margin. We propose an exploratory four-dimension Benchmark Maturity Index (BMI), scoring each capability domain on standardization, participation, independent verification, and national coverage. We show that the BMI refines, and in some cases revises, the maturity judgments that a purely descriptive review would produce. We argue that many apparent capability gaps in the public record cannot be distinguished, on available evidence, from evaluation-ecosystem gaps. This has direct implications for how national AI programs should design monitoring and funding criteria.
Summary
Main Finding
Publicly reported benchmark results show Indian foundation models demonstrate strong performance on older, now-saturated benchmarks (e.g., MMLU, MATH-500, HumanEval/MBPP). However, Indian models participate far less in newer, agentic, and domain-specialized evaluations that better discriminate frontier capability. Benchmark participation and disclosure are uneven across Indian organizations (Sarvam AI has the broadest public coverage). The paper introduces a Benchmark Maturity Index (BMI) to distinguish true capability gaps from gaps in the public evaluation ecosystem — a distinction that matters for monitoring, funding, and national AI policy.
Key Points
- Benchmark snapshot (August 2026): Indian models score highly on several established benchmarks but those benchmarks are widely regarded as saturated at the frontier.
- Frontier developers increasingly stop reporting on saturated benchmarks, shifting attention to harder, agentic, or domain-specific evaluations; Indian models rarely appear on these newer benchmarks.
- Benchmark disclosure is highly uneven among IndiaAI-supported organizations: only Sarvam AI provides broad multi-domain public results; many supported groups publish little or no standardized benchmark evidence.
- The paper proposes a three-tier working definition of "Indian-developed" models:
- Tier 1: fully indigenous (trained from scratch on India-controlled compute and significant India-sourced data).
- Tier 2: India-led with global components.
- Tier 3: India-adapted foreign base models (fine-tuned/adapted).
- Introduces a four-dimension Benchmark Maturity Index (BMI) at the capability-domain level, scoring:
- Standardization (existence of accepted benchmarks),
- Participation (number/representativeness of models evaluated),
- Independent verification (third-party/leaderboard confirmation),
- National coverage (domestic developer participation and disclosure).
- The BMI refines and sometimes revises maturity judgments from qualitative review alone; many apparent capability gaps are equally plausibly explained by evaluation-ecosystem gaps.
- Key caveat: analysis uses only publicly reported scores (developer-reported vs independently verified are distinguished) and is therefore a snapshot of the public record, not the full state of capability.
Data & Methods
- Model groups compared:
- Frontier: GPT-5.6, Claude Opus 5, Kimi K3, Qwen3.8-Max (top models on third-party leaderboards with public technical reports).
- Comparable-scale: models defined by active-parameter range (~12–50B active): Inkling (MoE), Qwen3.6-27B, Nemotron 3 Super.
- Indian models: Sarvam-105B, Sarvam-30B, Param2 (BharatGen), Krutrim-2; plus domain-specific Indian systems (e.g., Shodh AI’s Project Skanda, Avataar AI’s Varya).
- Capability domains (eight): general-purpose reasoning/knowledge, coding/software engineering, agentic AI/computer use, cybersecurity, vision/image understanding, video/multimodal understanding, scientific research, Indic language capability.
- Benchmarks used (examples): MMLU, GPQA Diamond, HLE, MATH-500, ARC-AGI; HumanEval, MBPP, LiveCodeBench, SWE-bench Pro, Terminal-Bench 2.1; domain-specific leaderboards where applicable.
- Data sources: only publicly available benchmark results from technical reports, model cards, and third‑party leaderboards. No independent evaluations or re-runs were performed by the authors.
- Distinctions made: developer-reported vs independently verified scores; BMI participation dimension scored against the four-model frontier sample (sensitivity noted).
- Timebound: data collection and analysis reflect the public record as of August 2026.
Implications for AI Economics
-
Measurement and inference
- Public benchmark scores are an imperfect proxy for national AI capability. Relying on public benchmarks risks under- or mis-estimating capability where disclosure is uneven.
- Benchmark-saturation effects mean high scores on older tests do not imply frontier competitiveness; economic assessments should weight benchmark discriminativity and recency.
- Evaluation-ecosystem maturity (availability, standardization, independent verification) must be modeled separately from model capability in national-level productivity and capability metrics.
-
Policy and funding design
- Monitoring and funding criteria for publicly supported model programs should require standardized benchmark disclosure and independent verification to avoid conflating reporting gaps with capability gaps.
- Differentiate funding by development tier: Tier 1 efforts may need sustained public investment in compute, data, and engineering capacity; Tier 2/3 work may require targeted incentives (e.g., fine-tuning/data curation).
- Support participation in harder, agentic, and domain-specific benchmarks (which better predict real-world economic impact) via grants, benchmarking credits, or centralized evaluation infrastructure.
-
Strategic and market effects
- Transparent public benchmarking is a signal to purchasers, partners, and labor markets; uneven disclosure may distort procurement decisions and private investment.
- National benchmarking infrastructure and standardized public leaderboards can increase the visibility and comparability of domestic models, reducing information asymmetries that affect market entry, procurement, and international collaboration.
- Investing in independent benchmarking capacity (national evaluation hubs, standardized suites for Indic languages and domain tasks) creates public goods that lower assessment costs for both public and private actors and improves the accuracy of economic indicators tied to AI capability.
-
Research and workforce development
- Emphasize benchmarks that reflect multi-step, agentic, and domain-specific tasks to better align evaluation with economic use-cases (software engineering, scientific discovery, multimodal services).
- Public benchmarking programs can stimulate skills development, standardize evaluation practice, and create datasets and tooling with spillovers across academia and industry.
Recommendations (concise) - Require public disclosure of standardized benchmark suites and encourage independent verification for recipients of public funds. - Invest in national benchmarking infrastructure and updated benchmark suites that include agentic and domain-specialized evaluations, plus strong Indic-language tests. - Use the BMI (or a refined version) as part of funding/monitoring dashboards to distinguish capability from evaluation-maturity gaps. - Tailor public funding by model tier: prioritize compute/data for Tier 1 and incentivize participation in advanced benchmarks for all tiers.
Limitations to bear in mind - The study uses only public reports (snapshot August 2026); non-disclosure does not imply lack of capability. - Active parameter counts and architecture heterogeneity limit strict comparability across models. - BMI is exploratory and intended as a framework for further development rather than a validated index.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The publicly benchmarked Indian models achieve strong scores on MMLU and MATH-500, but these benchmarks are saturated and do not establish frontier competitiveness. Output Quality | mixed | Benchmark performance on general-purpose knowledge and mathematical reasoning |
Reading fidelity
high
Study strength
medium
|
n=11
|
| On GPQA Diamond, Sarvam-105B scores 78.7, trailing the surveyed frontier leader by roughly 15 points and the best comparable-scale model by 8.5 points. Output Quality | negative | General-purpose reasoning performance on GPQA Diamond |
Reading fidelity
high
Study strength
medium
|
n=11
roughly 15 points behind the frontier leader; 8.5 points behind the comparable-scale leader
|
| On HLE with tools, the best comparable-scale model substantially outperforms Sarvam-105B, scoring 46.0 versus 11.2. Output Quality | negative | Hard reasoning performance on Humanity's Last Exam with tools |
Reading fidelity
high
Study strength
medium
|
n=11
34.8-point score difference
|
| No Indian model reports a score on ARC-AGI-2 or ARC-AGI-3. Adoption Rate | null_result | Public benchmark participation in advanced reasoning evaluations |
Reading fidelity
high
Study strength
low
|
n=4
|
| Indian models report strong scores on foundational coding benchmarks, including HumanEval, MBPP, and LiveCodeBench. Developer Productivity | positive | Foundational code-generation benchmark performance |
Reading fidelity
high
Study strength
medium
|
n=4
92.1 on HumanEval; 92.7 on MBPP; 71.7 on LiveCodeBench
|
| No Indian model reports a result on any of the four agentic or multi-step software-engineering benchmarks reviewed. Adoption Rate | null_result | Public participation in agentic and multi-step software-engineering evaluations |
Reading fidelity
high
Study strength
low
|
n=4
|
| The available public evidence shows Indian benchmark participation concentrated in foundational coding, while it remains absent from the reviewed agentic coding benchmarks. Adoption Rate | mixed | Distribution of public benchmark participation across coding-evaluation types |
Reading fidelity
high
Study strength
low
|
n=4
|
| The paper cannot determine from available evidence whether the absence of Indian results on agentic coding benchmarks reflects a genuine capability gap or only a reporting gap. Ai Safety And Ethics | mixed | Interpretability of public benchmark nonparticipation as capability versus disclosure |
Reading fidelity
high
Study strength
low
|
n=4
|
| Among the 12 organizations supported by the IndiaAI Innovation Centre's Foundation Models pillar, only Sarvam AI publishes benchmark results across multiple capability domains; BharatGen/Param2 reports limited results, and the remaining 10 have not disclosed standardized scores within the assessed domains. Adoption Rate | negative | Public benchmark disclosure and evaluation coverage among supported organizations |
Reading fidelity
high
Study strength
medium
|
n=12
1 of 12 organizations with broad multi-domain benchmark coverage; 10 of 12 with no disclosed standardized scores in assessed domains
|
| Sarvam AI reports the broadest benchmark coverage among the surveyed Indian organizations by a substantial margin. Adoption Rate | positive | Breadth of public benchmark participation across capability domains |
Reading fidelity
high
Study strength
medium
|
n=12
|
| The proposed Benchmark Maturity Index refines, and in two cases revises, maturity judgments produced by a purely qualitative review. Governance And Regulation | positive | Maturity assessment of the national AI evaluation ecosystem |
Reading fidelity
high
Study strength
low
|
n=8
two cases revised
|