The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A new expert-designed benchmark shows top AI models can complete high-value knowledge tasks at roughly 60–64% of expert-level quality, with GPT‑5 leading; however, the sizable gap to human professionals underscores continued limits to AI replacing high-skilled knowledge work.

The AI Productivity Index (APEX)
Bertie Vidgen, Abby Fennelly, Evan Pinnix, Chirag Mahapatra, Zach Richards, Austin Bridges, Calix Huang, Ben Hunsberger, Fez Zafar, Brendan Foody, Dominic Barton, Cass R. Sunstein, Eric Topol, Osvald Nitski · January 22, 2026 · SuperIntelligence - Robotics - Safety & Alignment
openalex descriptive medium evidence 8/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Bertie Vidgen provider ID
  2. Abby Fennelly provider ID
  3. Evan Pinnix provider ID
  4. Chirag Mahapatra provider ID
  5. Zach Richards provider ID
  6. Austin Bridges provider ID
  7. Calix Huang provider ID
  8. Ben Hunsberger provider ID
  9. Fez Zafar provider ID
  10. Brendan Foody provider ID
  11. Dominic Barton provider ID
  12. Cass R. Sunstein provider ID
  13. Eric Topol provider ID
  14. Osvald Nitski provider ID

Semantic Scholar

Latest observation:

  1. Bertie Vidgen provider ID
  2. Abby Fennelly provider ID
  3. Evan Pinnix provider ID
  4. Chirag Mahapatra provider ID
  5. Zach Richards provider ID
  6. Austin Bridges provider ID
  7. Calix Huang provider ID
  8. Ben Hunsberger provider ID
  9. Fez Zafar provider ID
  10. Brendan Foody provider ID
  11. Dominic Barton provider ID
  12. C. Sunstein provider ID
  13. Eric Topol provider ID
  14. Osvald Nitski provider ID
An expert-built benchmark (APEX-v1.0) finds leading frontier models score around 60–64% on high-value knowledge-work tasks across banking, consulting, law, and primary care, but a substantial gap remains relative to human experts.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We introduce the first version of the AI Productivity Index (APEX), a benchmark for assessing whether frontier AI models can perform knowledge work with high economic value. APEX addresses one of the largest inefficiencies in AI research: outside of coding, benchmarks often fail to test economically relevant capabilities. APEX-v1.0 contains 200 test cases and covers four domains: investment banking, management consulting, law, and primary medical care. It was built in three steps. First, we sourced experts with top-tier experience e.g., investment bankers from Goldman Sachs. Second, experts created prompts that reflect high-value tasks in their day-to-day work. Third, experts created rubrics for evaluating model responses. We evaluate 23 frontier models on APEX-v1.0 using an LM judge. GPT 5 (Thinking = High) achieves the highest mean score (64.2%), followed by Grok 4 (61.3%) and Gemini 2.5 Flash (Thinking = On) (60.4%). Qwen 3 235B is the best performing opensource model and seventh best overall. There is a large gap between the performance of even the best models and human experts, highlighting the need for better measurement of models’ ability to produce economically valuable work.

Summary

Main Finding

APEX-v1.0 is a new, expert-designed benchmark that evaluates whether frontier LMs can perform high-value knowledge-work tasks with economic relevance. On the 200-case held-out suite (investment banking, management consulting, law, primary care), the best models achieve roughly 60–64% of rubric criteria on average (GPT‑5: 64.2%, Grok 4: 61.3%, Gemini 2.5 Flash: 60.4%), with a substantial gap to human expert performance. This indicates strong but incomplete capability for producing economically valuable outputs and highlights the need for task-aligned, economically grounded measurement of AI progress.

Key Points

  • Purpose: APEX measures models on realistic, high-value knowledge-work tasks (1–8 hours of expert time, mean 3.5 hours), rather than abstract capabilities.
  • Dataset:
    • 200 held-out test cases, evenly split across four domains: investment banking, management consulting, law, primary care (medicine).
    • 76 domain experts contributed; mean experience ≈ 7.25 years.
    • 5,818 rubric criteria total; mean ≈ 29 criteria per case (range 7–54).
    • Mean evidence sources per case ≈ 5.8; mean total source tokens per case ≈ 26,677 (max limited to ~100k tokens).
  • Evaluation procedure:
    • 23 frontier models evaluated (13 closed-source, 10 open-source).
    • Each model run 3× per case; median score reported.
    • Responses autograded by a 3-LM judge panel (o3 low, Gemini 2.5 Pro off, Sonnet 4 off) with majority-vote Pass/Fail per rubric criterion.
    • Judge panel shows high internal consistency and ~89% agreement with human labels (on one model’s annotations).
  • Summary results:
    • Top mean scores: GPT‑5 (64.2%), Grok 4 (61.3%), Gemini 2.5 Flash (60.4%).
    • Domain-average scores across models: medicine 47.5%, investment banking 47.6%, management consulting 52.6%, law 56.9%.
    • Large variation at the bottom of leaderboard; several models <50%.
  • Reliability & limitations noted by authors:
    • LM judges can be biased; using a panel mitigates but does not eliminate bias.
    • Judges used are themselves models (with different “Thinking” settings), which required calibration checks (self-preference and inter-judge agreement).
    • Responses and model behavior are non-deterministic (mean range across 3 runs ≈ 11.9 percentage points).
    • APEX-v1.0 is a closed held-out dataset to preserve rigorous evaluation.

Data & Methods

  • Expert sourcing and prompt/rubric creation:
    • Experts recruited and vetted (30–45 min interviews + 1–2 hour assessments); contributors produced prompts, evidence, and detailed binary rubrics decomposing “quality” into objective criteria (analogous to unit tests).
    • Prompts grounded in common high-value tasks for senior roles at top firms (e.g., Goldman Sachs, McKinsey, BigLaw, top medical centers).
    • Quality control: multi-stage human review and LM-assisted feedback; 300 prompts started, 200 accepted.
  • Dataset characteristics:
    • Mean prompt tokens ≈ 430 (domain variation: medicine shorter, consulting longer).
    • Rubrics allow per-criterion Pass/Fail → percentage-of-criteria-passed scalar score per response.
  • Models & inference settings:
    • 23 models (mostly 2025 releases); used recommended temperatures and “Thinking” (chain-of-thought) settings where available; no uniform system prompt except for Nova Pro.
    • Context windows capped to ensure sources fit across models.
  • Grading:
    • Panel of three judge LMs grades each criterion independently; majority vote decides Pass/Fail.
    • Evaluated judge performance on consistency, inter-judge agreement (3/3 agreement ≈ 81%), judge self-preference (small biases), and judge-human agreement (≈ 89% on one model’s human labels).
  • Metrics reported:
    • Mean score (percentage of rubric criteria passed).
    • Pairwise head-to-head win rates across tasks.
    • Frequency of being ranked first / last across cases.
    • Domain-specific mean scores.

Implications for AI Economics

  • Measurement alignment: APEX demonstrates the value of task-aligned benchmarks for estimating economic impact. Abstract capability tests (e.g., general reasoning or NLP metrics) can miss whether outputs are actually useful for high-paid knowledge work.
  • Productivity and augmentation vs replacement:
    • Best models reach ~60–64% of rubric criteria on average, indicating meaningful capability to augment expert workflows (drafting, triage, first-pass analysis), but not yet reliable autonomous replacement for complex expert tasks.
    • Heterogeneity by domain suggests partial automation potential: legal and consulting tasks appear easier on average than primary-care medical tasks (per rubric pass rates), so sectoral effects on labor demand will vary.
  • Labor market and firm strategy:
    • Firms can likely deploy LMs for assisted workflows (efficiency gains, faster first drafts, decision support), but must retain human oversight for correctness, liability, and nuanced judgment.
    • Investment in complementary human capital (supervision, validation, prompt engineering, rubric design) will be economically valuable.
  • Macroeconomic impact: Given the gap to human-level performance, near-term contributions to GDP via direct task automation may be limited and concentrated — large productivity gains are conditional on further model improvements, cost-effective deployment, and integration with workflows.
  • Policy and governance:
    • Benchmarks like APEX enable regulators and policymakers to assess sector-specific AI readiness and risk (e.g., medical/law liability, misinformation, malpractice).
    • Standardized, economically meaningful evaluation supports targeted workforce transition policies and standards for deployment (certification, audit trails, evaluation thresholds).
  • Research and investment priorities:
    • Need for broader, open, and domain-diverse benchmarks linked to economic outcomes (billing rates, time saved, error costs).
    • Improve human-ground-truth evaluations and judge calibration (reduce LM-judge bias).
    • Study cost-performance tradeoffs: model inference costs per task vs. economic value generated.
    • Track downstream outcomes (hiring, wages, firm productivity) tied to deploying models that pass APEX-style tasks.

Takeaway: APEX provides an actionable, expert-grounded way to quantify how much frontier LMs can do economically valuable knowledge work. Current frontier models show substantial capability but meaningful gaps remain; the benchmark should inform firm adoption strategies, labor-market modeling, and policy design while motivating further work in task-aligned evaluation and judge calibration.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Provides systematic, expert-driven measurement of model performance on 200 high-value tasks across four knowledge-work domains and evaluates 23 frontier models, so the descriptive evidence about relative model performance is credible; however, conclusions about economic value or real-world productivity are limited by the benchmark's scope (4 domains, 200 cases), reliance on an LM judge for scoring, potential prompt/rubric bias from a non-representative expert set, and possible dataset leakage into model training. Methods Rigormedium — Design steps (recruiting senior-domain experts, using expert-generated prompts and rubrics, testing many models) indicate careful construction, but reliance on an automated LM judge (rather than independent human raters or multi-annotator consensus), limited sample size per domain, unclear inter-rater reliability, and potential for prompt or rubric alignment with specific models reduce methodological rigor. SampleAPEX-v1.0 comprises 200 test cases drawn from four domains (investment banking, management consulting, law, primary medical care); prompts and evaluation rubrics were authored by domain experts with top-tier experience (e.g., Goldman Sachs bankers); 23 frontier language models were evaluated (including GPT-5, Grok 4, Gemini 2.5 Flash, Qwen 3 235B), and scoring was conducted using an LM judge; comparisons against human expert performance are reported qualitatively (gap noted), though human baseline scoring details are not fully described in the summary. Themesproductivity human_ai_collab adoption GeneralizabilityOnly four knowledge-work domains included — results may not generalize to other professions or tasks (e.g., engineering, education)., 200 test cases provides limited coverage of within-domain task heterogeneity and edge cases., Experts sourced from top-tier firms may produce prompts/rubrics that do not reflect average or routine work, biasing difficulty or relevance., Scoring via an LM judge risks systematic bias or inconsistency compared with multi-human raters., Potential data leakage: models may have seen similar prompts or solutions during training, inflating performance., Geographic, linguistic, and regulatory context constraints (e.g., medical and legal tasks may be jurisdiction-specific)., Benchmarks measure isolated prompts, not end-to-end integration into workplace workflows or productivity gains in situ.

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We introduce the first version of the AI Productivity Index (APEX), a benchmark for assessing whether frontier AI models can perform knowledge work with high economic value. Other positive existence and scope of APEX-v1.0 benchmark
Reading fidelity high
Study strength high
not reported
0.3
APEX addresses one of the largest inefficiencies in AI research: outside of coding, benchmarks often fail to test economically relevant capabilities. Other negative coverage of economically relevant capabilities by benchmarks
Reading fidelity high
Study strength speculative
not reported
0.03
APEX-v1.0 contains 200 test cases and covers four domains: investment banking, management consulting, law, and primary medical care. Other positive benchmark scope (number of test cases and domain coverage)
Reading fidelity high
Study strength high
n=200
0.3
APEX-v1.0 was built in three steps: (1) sourced experts with top-tier experience (e.g., investment bankers from Goldman Sachs), (2) experts created prompts reflecting high-value tasks in their day-to-day work, (3) experts created rubrics for evaluating model responses. Other positive benchmark construction process and provenance of prompts/rubrics
Reading fidelity high
Study strength high
not reported
0.3
We evaluate 23 frontier models on APEX-v1.0 using an LM judge. Other neutral model performance evaluations on APEX-v1.0
Reading fidelity high
Study strength high
n=23
0.3
GPT 5 (Thinking = High) achieves the highest mean score (64.2%), followed by Grok 4 (61.3%) and Gemini 2.5 Flash (Thinking = On) (60.4%). Output Quality positive mean APEX score (model output quality on benchmark)
Reading fidelity high
Study strength high
n=200
64.2%, 61.3%, 60.4%
0.3
Qwen 3 235B is the best performing opensource model and seventh best overall. Output Quality positive model ranking among evaluated models on APEX-v1.0
Reading fidelity high
Study strength high
n=23
7th place (best open-source)
0.3
There is a large gap between the performance of even the best models and human experts, highlighting the need for better measurement of models’ ability to produce economically valuable work. Output Quality negative gap between model performance and human expert performance on economically valuable tasks
Reading fidelity medium
Study strength medium
not reported
0.11

Notes