0 cumulative citations
View corpus contextA new expert-designed benchmark shows top AI models can complete high-value knowledge tasks at roughly 60–64% of expert-level quality, with GPT‑5 leading; however, the sizable gap to human professionals underscores continued limits to AI replacing high-skilled knowledge work.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextWe introduce the first version of the AI Productivity Index (APEX), a benchmark for assessing whether frontier AI models can perform knowledge work with high economic value. APEX addresses one of the largest inefficiencies in AI research: outside of coding, benchmarks often fail to test economically relevant capabilities. APEX-v1.0 contains 200 test cases and covers four domains: investment banking, management consulting, law, and primary medical care. It was built in three steps. First, we sourced experts with top-tier experience e.g., investment bankers from Goldman Sachs. Second, experts created prompts that reflect high-value tasks in their day-to-day work. Third, experts created rubrics for evaluating model responses. We evaluate 23 frontier models on APEX-v1.0 using an LM judge. GPT 5 (Thinking = High) achieves the highest mean score (64.2%), followed by Grok 4 (61.3%) and Gemini 2.5 Flash (Thinking = On) (60.4%). Qwen 3 235B is the best performing opensource model and seventh best overall. There is a large gap between the performance of even the best models and human experts, highlighting the need for better measurement of models’ ability to produce economically valuable work.
Summary
Main Finding
APEX-v1.0 is a new, expert-designed benchmark that evaluates whether frontier LMs can perform high-value knowledge-work tasks with economic relevance. On the 200-case held-out suite (investment banking, management consulting, law, primary care), the best models achieve roughly 60–64% of rubric criteria on average (GPT‑5: 64.2%, Grok 4: 61.3%, Gemini 2.5 Flash: 60.4%), with a substantial gap to human expert performance. This indicates strong but incomplete capability for producing economically valuable outputs and highlights the need for task-aligned, economically grounded measurement of AI progress.
Key Points
- Purpose: APEX measures models on realistic, high-value knowledge-work tasks (1–8 hours of expert time, mean 3.5 hours), rather than abstract capabilities.
- Dataset:
- 200 held-out test cases, evenly split across four domains: investment banking, management consulting, law, primary care (medicine).
- 76 domain experts contributed; mean experience ≈ 7.25 years.
- 5,818 rubric criteria total; mean ≈ 29 criteria per case (range 7–54).
- Mean evidence sources per case ≈ 5.8; mean total source tokens per case ≈ 26,677 (max limited to ~100k tokens).
- Evaluation procedure:
- 23 frontier models evaluated (13 closed-source, 10 open-source).
- Each model run 3× per case; median score reported.
- Responses autograded by a 3-LM judge panel (o3 low, Gemini 2.5 Pro off, Sonnet 4 off) with majority-vote Pass/Fail per rubric criterion.
- Judge panel shows high internal consistency and ~89% agreement with human labels (on one model’s annotations).
- Summary results:
- Top mean scores: GPT‑5 (64.2%), Grok 4 (61.3%), Gemini 2.5 Flash (60.4%).
- Domain-average scores across models: medicine 47.5%, investment banking 47.6%, management consulting 52.6%, law 56.9%.
- Large variation at the bottom of leaderboard; several models <50%.
- Reliability & limitations noted by authors:
- LM judges can be biased; using a panel mitigates but does not eliminate bias.
- Judges used are themselves models (with different “Thinking” settings), which required calibration checks (self-preference and inter-judge agreement).
- Responses and model behavior are non-deterministic (mean range across 3 runs ≈ 11.9 percentage points).
- APEX-v1.0 is a closed held-out dataset to preserve rigorous evaluation.
Data & Methods
- Expert sourcing and prompt/rubric creation:
- Experts recruited and vetted (30–45 min interviews + 1–2 hour assessments); contributors produced prompts, evidence, and detailed binary rubrics decomposing “quality” into objective criteria (analogous to unit tests).
- Prompts grounded in common high-value tasks for senior roles at top firms (e.g., Goldman Sachs, McKinsey, BigLaw, top medical centers).
- Quality control: multi-stage human review and LM-assisted feedback; 300 prompts started, 200 accepted.
- Dataset characteristics:
- Mean prompt tokens ≈ 430 (domain variation: medicine shorter, consulting longer).
- Rubrics allow per-criterion Pass/Fail → percentage-of-criteria-passed scalar score per response.
- Models & inference settings:
- 23 models (mostly 2025 releases); used recommended temperatures and “Thinking” (chain-of-thought) settings where available; no uniform system prompt except for Nova Pro.
- Context windows capped to ensure sources fit across models.
- Grading:
- Panel of three judge LMs grades each criterion independently; majority vote decides Pass/Fail.
- Evaluated judge performance on consistency, inter-judge agreement (3/3 agreement ≈ 81%), judge self-preference (small biases), and judge-human agreement (≈ 89% on one model’s human labels).
- Metrics reported:
- Mean score (percentage of rubric criteria passed).
- Pairwise head-to-head win rates across tasks.
- Frequency of being ranked first / last across cases.
- Domain-specific mean scores.
Implications for AI Economics
- Measurement alignment: APEX demonstrates the value of task-aligned benchmarks for estimating economic impact. Abstract capability tests (e.g., general reasoning or NLP metrics) can miss whether outputs are actually useful for high-paid knowledge work.
- Productivity and augmentation vs replacement:
- Best models reach ~60–64% of rubric criteria on average, indicating meaningful capability to augment expert workflows (drafting, triage, first-pass analysis), but not yet reliable autonomous replacement for complex expert tasks.
- Heterogeneity by domain suggests partial automation potential: legal and consulting tasks appear easier on average than primary-care medical tasks (per rubric pass rates), so sectoral effects on labor demand will vary.
- Labor market and firm strategy:
- Firms can likely deploy LMs for assisted workflows (efficiency gains, faster first drafts, decision support), but must retain human oversight for correctness, liability, and nuanced judgment.
- Investment in complementary human capital (supervision, validation, prompt engineering, rubric design) will be economically valuable.
- Macroeconomic impact: Given the gap to human-level performance, near-term contributions to GDP via direct task automation may be limited and concentrated — large productivity gains are conditional on further model improvements, cost-effective deployment, and integration with workflows.
- Policy and governance:
- Benchmarks like APEX enable regulators and policymakers to assess sector-specific AI readiness and risk (e.g., medical/law liability, misinformation, malpractice).
- Standardized, economically meaningful evaluation supports targeted workforce transition policies and standards for deployment (certification, audit trails, evaluation thresholds).
- Research and investment priorities:
- Need for broader, open, and domain-diverse benchmarks linked to economic outcomes (billing rates, time saved, error costs).
- Improve human-ground-truth evaluations and judge calibration (reduce LM-judge bias).
- Study cost-performance tradeoffs: model inference costs per task vs. economic value generated.
- Track downstream outcomes (hiring, wages, firm productivity) tied to deploying models that pass APEX-style tasks.
Takeaway: APEX provides an actionable, expert-grounded way to quantify how much frontier LMs can do economically valuable knowledge work. Current frontier models show substantial capability but meaningful gaps remain; the benchmark should inform firm adoption strategies, labor-market modeling, and policy design while motivating further work in task-aligned evaluation and judge calibration.
Assessment
Claims (8)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We introduce the first version of the AI Productivity Index (APEX), a benchmark for assessing whether frontier AI models can perform knowledge work with high economic value. Other | positive | existence and scope of APEX-v1.0 benchmark |
Reading fidelity
high
Study strength
high
|
not reported
|
| APEX addresses one of the largest inefficiencies in AI research: outside of coding, benchmarks often fail to test economically relevant capabilities. Other | negative | coverage of economically relevant capabilities by benchmarks |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| APEX-v1.0 contains 200 test cases and covers four domains: investment banking, management consulting, law, and primary medical care. Other | positive | benchmark scope (number of test cases and domain coverage) |
Reading fidelity
high
Study strength
high
|
n=200
|
| APEX-v1.0 was built in three steps: (1) sourced experts with top-tier experience (e.g., investment bankers from Goldman Sachs), (2) experts created prompts reflecting high-value tasks in their day-to-day work, (3) experts created rubrics for evaluating model responses. Other | positive | benchmark construction process and provenance of prompts/rubrics |
Reading fidelity
high
Study strength
high
|
not reported
|
| We evaluate 23 frontier models on APEX-v1.0 using an LM judge. Other | neutral | model performance evaluations on APEX-v1.0 |
Reading fidelity
high
Study strength
high
|
n=23
|
| GPT 5 (Thinking = High) achieves the highest mean score (64.2%), followed by Grok 4 (61.3%) and Gemini 2.5 Flash (Thinking = On) (60.4%). Output Quality | positive | mean APEX score (model output quality on benchmark) |
Reading fidelity
high
Study strength
high
|
n=200
64.2%, 61.3%, 60.4%
|
| Qwen 3 235B is the best performing opensource model and seventh best overall. Output Quality | positive | model ranking among evaluated models on APEX-v1.0 |
Reading fidelity
high
Study strength
high
|
n=23
7th place (best open-source)
|
| There is a large gap between the performance of even the best models and human experts, highlighting the need for better measurement of models’ ability to produce economically valuable work. Output Quality | negative | gap between model performance and human expert performance on economically valuable tasks |
Reading fidelity
medium
Study strength
medium
|
not reported
|