The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models already perform strongly on business-school case analyses, scoring highly against instructor rubrics and improving sharply over two years; while this suggests rapid gains on complex analytical tasks that underpin entry-level professional work, benchmark results may overstate real-world impact due to rubric subjectivity and potential data overlap.

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani, Mitch Weiss · July 17, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Semantic Scholar

Latest observation:

  1. Ajay Patel provider ID
  2. K. Hosanagar provider ID
  3. Ramayya Krishnan provider ID
  4. Christopher Callison-Burch provider ID
  5. Karim R. Lakhani provider ID
BusinessCaseBench—an instructor-rubric-scored benchmark of hundreds of business-school case questions—shows frontier LLMs already score highly and have improved substantially over two years on tasks resembling analytical white-collar work.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.

Summary

Main Finding

Frontier LLMs already perform at a high level on open-ended, expert-written business case work across 18 business disciplines. When graded against instructor-derived rubrics, leading models achieve ≈81–88% average coverage of expected rubric elements (partial-credit Standard scoring), though full rubric satisfaction (Complete Answer scoring) is much lower (≈32–50%). Model capability has improved rapidly: a within-family comparison documents roughly a 23-percentage-point gain over ~2 years.

Key Points

  • Benchmark introduced: BusinessCaseBench — 615 open-ended questions from 238 business school cases spanning 18 disciplines, each paired with instructor case solutions transformed into checklist rubrics.
  • Models evaluated: OpenAI GPT-5.4, Anthropic Claude Sonnet 4.6, Google Gemini 3 Flash Preview; plus a four-model within-family (OpenAI) time series spanning ~2 years.
  • Two scoring metrics:
    • Standard (partial credit): rubric-weighted fraction of checklist items satisfied.
    • Complete Answer (all-or-nothing): only counts if every checklist item is satisfied.
  • Aggregate performance:
    • Standard scores (mean): Claude Sonnet 4.6 88.4%, GPT-5.4 87.2%, Gemini 3 Flash 81.6%.
    • Complete Answer: Claude 49.6%, GPT-5.4 47.6%, Gemini 32.0%.
  • Discipline variation exceeds model-provider gaps: Standard scores by discipline ranged roughly 80.1% (Marketing & Sales) to 95.0% (Business & Government Relations). Complete Answer range wider and lower.
  • Most questions are partially addressed well; fully complete answers are less common — AI outputs often look like high-quality drafts requiring human review.
  • Less than 7% of questions were missed by all frontier models, implying required knowledge is present across models and ceilings are rare.
  • Coarse metadata (numerical vs non-numerical, subjective vs objective, fictional vs real cases) explain little of the score variation; most difficulty is case- and question-specific.
  • Validation: an LLM-as-judge protocol scores model outputs against rubrics, with human annotator checks on a subset to validate automated grading.

Data & Methods

  • Data composition:
    • 238 instructor-written business cases → 615 questions across 18 disciplines.
    • Labels: fictional vs real case, numerical vs non-numerical, subjective vs objective.
    • Mapped to O*NET taxonomy (24 Work Activities, 55 Intermediate WAs, 108 Detailed WAs) so results can be tied to occupational task taxonomy.
  • Rubric construction:
    • Instructor case solutions parsed into equally-weighted checklist rubrics (each checklist item is a required analytic element).
  • Evaluation pipeline:
    • Provide full case narrative + question prompt to model → model produces open-ended answer.
    • Use an LLM-as-judge to compare model answer to checklist rubric items and assign scores (Standard and Complete Answer).
    • Bootstrapped 95% CIs reported; model identifiers and sampling settings pinned.
    • Human validation: for a subset, three trained human annotators independently authored rubrics and graded blinded model responses to check alignment with automated scoring.
  • Models & comparisons:
    • Cross-provider snapshot: GPT-5.4, Claude Sonnet 4.6, Gemini 3 Flash Preview (all 615 questions).
    • Longitudinal: four OpenAI model generations (e.g., GPT-4 Turbo → GPT-5.4) on same benchmark to quantify generational improvement (~23 percentage points gain).
  • Analyses:
    • Regression and cross-validation to partition explained variance by discipline and question-type strata (both explain little of total variance; most variation is question-specific).
    • Rank concordance across models high → discipline-level difficulty consistent across providers.

Implications for AI Economics

  • Task-level substitution and augmentation
    • High partial-credit performance on open-ended analytical work implies AI can produce analytically useful drafts for many white-collar tasks historically performed by entry-level analysts, consultants, and managers.
    • Because complete, defensible answers remain less common, the most immediate impact is likely task augmentation (AI drafting + human review) rather than wholesale replacement for complex judgment tasks.
  • Occupational reallocation
    • Mapping to O*NET allows identification of specific work activities at risk (structured analytical/explanatory activities approach higher ceilings; open-ended advisory, opportunity-identification, and multi-stakeholder judgment tasks remain harder).
    • Rapid capability gains (23 pp in ~2 years within one model family) imply short horizons for substantial reallocation in junior professional roles; firms should anticipate shifting task bundles and reconfigure job designs.
  • Productivity and wage effects
    • AI-generated analytical drafts can raise per-worker productivity for tasks with routinizable checklist components, potentially compressing demand for raw execution and increasing demand for supervision, judgment, synthesis, and client-facing skills.
    • Distributional effects: entry-level, high-volume analytical work faces downward pressure; premium placed on tasks where humans add comparative advantage (persuasion, negotiation, contextual judgment).
  • Education and human capital policy
    • Business-school case pedagogy trains precisely the tasks measured; given high AI capability on these tasks, curricula should pivot toward skills complementary to AI (strategic leadership, interpersonal negotiation, ethical judgment, execution in uncertainty, critical review of AI outputs).
    • Training programs should emphasize human-in-the-loop skills: prompt engineering for domain relevance, verification, error-detection, and high-level strategy.
  • Organizational design & governance
    • Firms should redesign workflows: treat LLM outputs as draft inputs requiring verification; establish review protocols, audit trails, and responsibility allocation to avoid overreliance and legal/ethical risk.
    • Because knowledge is distributed across models and fewer than 7% of questions are universally unsolved, ensembles, retrieval augmentation, or tool-chaining could materially raise performance quickly.
  • Policy considerations
    • Labor-market monitoring: track occupation- and task-level displacement risk using O*NET mappings; support re-skilling targeted at tasks where humans retain comparative advantage.
    • Assessment & credentialing: as AI becomes capable of producing high-quality analytic drafts, evaluation systems (education, professional credentialing, hiring) should distinguish between human-authored judgment and AI-assisted outputs.
    • Risk management: ensure standards for transparency, bias/fairness checks, and liability assignment as AI is embedded into decision workflows.
  • Research & measurement
    • BusinessCaseBench provides a task-focused, discipline-spanning instrument that shifts the empirical question from "can AI do this work?" to "how will task allocation and institutions adapt as capability redistributes?" — enabling more granular economic analysis of diffusion, productivity, and labor-market impacts.

Concise recommendations for stakeholders: - Firms: pilot AI-assistants for drafting analytic work, reallocate human effort to review and stakeholder-facing tasks, and update job specs. - Educators: teach critical evaluation of AI outputs, negotiation/leadership, and domain judgment; integrate AI tools into pedagogy as assistants rather than substitutes. - Policymakers: fund targeted retraining, update occupational statistics to task-level measurement, and mandate governance standards for high-stakes decision use.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides systematic, quantitative measurement of LLM performance on a novel task set (business-school case analyses) and documents sizable improvement over time, which is strong descriptive evidence about model capabilities; however, benchmark scores against instructor rubrics do not establish real-world effectiveness or causal impacts on productivity and are vulnerable to rubric alignment, subjectivity, and potential training-data leakage. Methods Rigormedium — The authors construct a multi-discipline dataset, derive grading rubrics from instructor solutions, and evaluate multiple frontier models longitudinally, which are sound methodological steps for a benchmark paper; limitations include potential rubric bias, limited detail on rubric validation and inter-rater reliability, possible overlap between benchmark content and model training data, and reliance on rubric-based scoring rather than external, real-world outcome measures. SampleA newly constructed benchmark (BusinessCaseBench) consisting of hundreds of questions drawn from business school case materials across 18 disciplines, each paired with an instructor-derived grading rubric; frontier LLM families are evaluated on these items, with comparisons showing substantial improvement within one model family over a two-year period. Themesproductivity skills_training human_ai_collab GeneralizabilityBenchmarks derived from business-school cases may not represent the full range of real-world white-collar analytical work (e.g., collaborative, iterative, or domain-specialist settings)., Instructor rubrics capture one notion of correctness/quality and may not align with workplace judgments or stakeholder trade-offs., Potential training-data leakage: models may have seen similar cases or instructor materials during pretraining, inflating scores., Subjective components of case work (judgment, persuasion, negotiation) are poorly captured by rubric scoring., Language and cultural bias if cases and rubrics are primarily English and Western business-school materials., Evaluation in isolated benchmark settings may not reflect performance under time pressure, multi-stakeholder interaction, or downstream implementation constraints.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Large language models (LLMs) are improving rapidly as reflected in benchmark scores. Other positive LLM benchmark performance (aggregate scores)
Reading fidelity high
Study strength medium
not reported
0.18
Existing AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. Other null_result scope/coverage of existing AI benchmarks
Reading fidelity high
Study strength medium
not reported
0.18
AI progress on analytical knowledge work performed by white-collar professionals — including synthesizing complex information, exercising judgment under uncertainty, applying strategic and adversarial thinking, weighing trade-offs, and producing defensible structured analyses — remains poorly measured by current benchmarks. Other null_result coverage of analytical knowledge-work skills in benchmarks
Reading fidelity high
Study strength low
not reported
0.09
The gap in measurement is even more pronounced for subjective components of analytical work, where success can be challenging to define. Other null_result measurability of subjective components of analytical work
Reading fidelity high
Study strength low
not reported
0.09
The 'case method' used by top business schools provides a natural foundation for addressing this measurement gap. Other positive suitability of case-method pedagogy for benchmarking analytical tasks
Reading fidelity high
Study strength low
not reported
0.09
We construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. Other null_result benchmark composition (number of questions and disciplinary coverage)
Reading fidelity high
Study strength high
not reported
0.3
On BusinessCaseBench, frontier AI models already score highly against instructor rubrics. Output Quality positive AI model performance vs. instructor rubrics on case questions
Reading fidelity high
Study strength medium
not reported
0.18
Capability within one model family improves substantially over two years. Output Quality positive change in model capability over time within a model family
Reading fidelity medium
Study strength medium
not reported
0.11
These results provide strong evidence that AI performance on this class of work is already high and rapidly improving. Output Quality positive AI performance on case-method analytical tasks
Reading fidelity high
Study strength medium
not reported
0.18
There are implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work. Employment mixed potential impacts on business school pedagogy and entry-level professional roles (education and employment implications)
Reading fidelity high
Study strength speculative
not reported
0.03

Notes