The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

State-of-the-art AI agents make substantial partial progress but complete fewer than one-third of real, market-validated workflows; shortcomings in complex instruction-following and domain expertise keep many professional deliverables out of reliable reach.

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang · August 18, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Liya Zhu unresolved corpus identity
  2. Xin Ma unresolved corpus identity
  3. Tao Liu unresolved corpus identity
  4. Haodong Wang unresolved corpus identity
  5. Ge Zhang unresolved corpus identity
  6. Jingzhe Ding unresolved corpus identity
  7. Qingshui Gu unresolved corpus identity
  8. Yongjie Zhong unresolved corpus identity
  9. Jinxiang Meng unresolved corpus identity
  10. Yuan Gao unresolved corpus identity
  11. Yunqiu Zhou unresolved corpus identity
  12. Hao Zhu unresolved corpus identity
  13. Jifeng He unresolved corpus identity
  14. Yongzhi Liao unresolved corpus identity
  15. Xinyi Zhang unresolved corpus identity
  16. Chaoxin Li unresolved corpus identity
  17. Yi Zhu unresolved corpus identity
  18. Xi Lin unresolved corpus identity
  19. Duju Zeng unresolved corpus identity
  20. Xiang Gao unresolved corpus identity
  21. Wen Zhang unresolved corpus identity
  22. Yunyang Wang unresolved corpus identity
  23. Duo Wang unresolved corpus identity
  24. Huan Zhou unresolved corpus identity
  25. Zuo Wang unresolved corpus identity
  26. Jin Chen unresolved corpus identity
  27. Kaiyuan Zhang unresolved corpus identity
  28. Chuqian Yu unresolved corpus identity
  29. Tianhao Yu unresolved corpus identity
  30. Longxiang Liu unresolved corpus identity
  31. Jianbo Xue unresolved corpus identity
  32. Huimin Che unresolved corpus identity
  33. Jiahao Wang unresolved corpus identity
  34. Yujia Qin unresolved corpus identity
  35. Jiaheng Liu unresolved corpus identity
  36. Shen Yan unresolved corpus identity
  37. Xiaolong Chang unresolved corpus identity
  38. Wenhao Huang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Liya Zhu provider ID
  2. Xin Ma provider ID
  3. Tao Liu provider ID
  4. Haodong Wang provider ID
  5. Ge Zhang provider ID
  6. Jingzhe Ding provider ID
  7. Qingshui Gu provider ID
  8. Yongjie Zhong provider ID
  9. Jinxiang Meng provider ID
  10. Yuan Gao provider ID
  11. Yun Zhou provider ID
  12. Hao Zhu provider ID
  13. Jifeng He provider ID
  14. Yong Liao provider ID
  15. Xinyi Zhang provider ID
  16. Chao Li provider ID
  17. Yi Zhu provider ID
  18. Xi Lin provider ID
  19. D. Zeng provider ID
  20. Xiang Gao provider ID
  21. Wen Zhang provider ID
  22. Yunyan Wang provider ID
  23. Duohongzi Wang provider ID
  24. Huan Zhou provider ID
  25. Zuo Wang provider ID
  26. Jin Chen provider ID
  27. Kaiyuan Zhang provider ID
  28. Chuqian Yu provider ID
  29. Tian Yu provider ID
  30. Longxiang Liu provider ID
  31. J. Xue provider ID
  32. Huimin Che provider ID
  33. Jiahao Wang provider ID
  34. Yujia Qin provider ID
  35. Jiaheng Liu provider ID
  36. Shen Yan provider ID
  37. Xiaolong Chang provider ID
  38. Wenhao Huang provider ID
Across 97 market-validated end-to-end startup workflows, state-of-the-art general-purpose agents achieve average task scores in the ~60–74% range but fully satisfy only roughly 30% of tasks, with failures concentrated in complex instruction-following and domain-specific expertise.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

Summary

Main Finding

StartupBench—an end-to-end benchmark built from market-validated AI startup workflows—shows that current general-purpose agent models make substantial partial progress on realistic professional tasks but reliably complete only a minority. Even the strongest evaluated models achieve average rubric-weighted scores in the low-to-mid 70s (out of 100) but succeed (score ≥ 90) on only ~30% or fewer tasks. Key failure modes include complex instruction following and gaps in domain-specific expertise, indicating that many commercially relevant workflows remain beyond reliable automation by generalist agents.

Project page: https://startupbench.github.io/ (paper dated Aug 19, 2026)

Key Points

  • Task sourcing: Tasks are derived from workflows in AI-native startups with demonstrated adoption (funding > USD 1M and paid/traction evidence), supplemented by 30+ user interviews and domain-expert reconstruction.
  • Benchmark scale & scope:
    • 97 end-to-end workflow tasks across 6 domains: Medical & Healthcare (21), Finance (18), Legal (16), Business & Management (19), STEM & CS (16), Education & Humanities (7).
    • Multi-format deliverables (DOCX, XLSX, PPTX, PDF, Markdown, images, code, etc.).
    • Rich evaluation: average 25.3 fine-grained rubric items per task, covering functionality, structure, formatting, and domain quality.
  • Task format: each task is a triple (q: user request, E: workspace/files, R: weighted rubric items). Evaluation aggregates binary rubric outcomes into a weighted score.
  • Evaluation methodology:
    • Agent-as-a-Judge: each rubric item judged by a lightweight judge agent (GPT-5.5 used as underlying judge) that inspects deliverables and evidence views.
    • Models run under a common agent harness (Nanobot), capped at 200 interaction steps.
    • 3 independent runs per model; 95% bootstrap CIs reported.
  • Models benchmarked: 9 representative models (closed- and open-source families), e.g., Kimi-K3, GPT-5.6-sol, GPT-5.5, Seed-2.1-Pro, Qwen-3.6-Max, DeepSeek-V4-Pro, GLM-5.1, etc.
  • Performance summary:
    • Best average scores: Kimi-K3 73.67, GPT-5.6-sol 73.61, GPT-5.5 72.79.
    • Success rate (score ≥ 90): top models ~26–31%; many domains (Finance, STEM, Education) are especially challenging.
  • Failure analysis: primary sources of failure were complex instruction following, nuanced domain expertise, and domain-specific constraints/formatting — not just basic language understanding or single-step reasoning.

Data & Methods

  • Data construction pipeline:
    • Startup survey to collect candidate agent products (20+ startups as seed pool).
    • User interviews (30+ deep users) to capture real workflows, inputs, outputs, constraints.
    • Domain experts (50+) converted interview-derived workflows into reproducible tasks and instances; tasks retained only if realistic, end-to-end, and evaluable.
    • Multi-stage quality control: expert cross-validation, pilot runs on frontier models (to calibrate difficulty and ensure discriminative value), rubric verification, and manual review.
  • Task filtering: tasks that were trivially solved by pilot models were removed to preserve discriminative power.
  • Evaluation harness:
    • Nanobot agent harness with uniform tool configuration; 200-step interaction cap per task.
    • Deliverable evidence view V(D): extracted text + rendered pages to aid judge agents.
    • AgentJudge per rubric: returns binary pass/fail and textual justification; final score is weighted aggregation of rubric outcomes.
  • Reliability measures: three independent runs per model and 10,000-sample bootstrap for 95% CIs.
  • Limitations noted by authors:
    • Startup selection bias (funding threshold → favors funded startups).
    • Judge agent uses GPT-5.5 which may introduce evaluator model bias.
    • Tasks were revised to be reproducible/evaluable which may prune some highly idiosyncratic real-world cases.
    • Evaluation limited to a unified agent harness and general-purpose agents; specialized vertical models/agents might perform differently.

Implications for AI Economics

  • Grounding benchmarks in market-validated products improves external validity for economic analysis:
    • StartupBench measures progress on tasks that users actually pay for, making it more informative for forecasts of automation, productization, and ROI.
  • Quantifying automation potential:
    • The gap between average rubric scores (~55–75 range) and low success rates (≤ ~30%) suggests substantial partial automation potential but limited full task automation today.
    • Economists can use the per-domain success rates to estimate asymmetric automation risk across sectors (e.g., Business & Management easier than Finance/STEM/Education).
  • Signals for investors and product strategy:
    • The benchmark highlights which professional workflows remain fragile for general-purpose agents—opportunity areas for vertical/specialized agents, domain-tuned models, and toolchain-integrated products.
    • Startups that tightly integrate domain expertise, deterministic tooling, and human-in-the-loop flows may capture higher-value use cases where generalists fail.
  • Labor market and productivity implications:
    • High partial-progress but low completion implies increased complementarities: agents can assist humans (e.g., drafts, analyses) but human oversight remains essential for final deliverables—supporting augmentation rather than immediate displacement in many roles.
    • Sectoral heterogeneity in task completion suggests uneven labor impacts; high-skill domain expertise tasks (Finance, STEM) are less automatable now, shifting near-term productivity gains toward workflows amenable to partial automation.
  • Valuation and adoption risk:
    • Benchmarked evidence may help investors/governance evaluate vendor claims; empirical, rubric-level performance is a useful signal to reduce technological overclaim risk.
    • Procurement decisions for enterprises can be better informed by end-to-end success rates and rubric-level failure modes (e.g., instruction following vs. factual knowledge).
  • Policy & regulation:
    • Demonstrated limitations in domain expertise and regulatory-constrained workflows (e.g., healthcare, legal) argue for caution in delegating critical decisions to generalist agents and for designing certification/validation regimes.
  • Research & market implications:
    • Indicates economic value in investing in domain-specialized models, structured tool access, better instruction-following, and evaluation pipelines that reflect market needs.
    • StartupBench itself can serve as an ongoing empirical indicator: improvements in mean scores and success rates over time can be mapped to likely pace of product-level automation and economic displacement/augmentation.

Potential uses for AI economists: - Map StartupBench success rates to estimates of task-level automation probabilities and sectoral exposure. - Use rubric-level failure modes to forecast which complementary investments (human training, verification tooling, domain data) yield highest productivity returns. - Evaluate startup technology claims and guide investment due diligence using E2E, market-grounded performance rather than isolated capability metrics.

If you want, I can: - Convert the benchmark results into a simple sector-by-sector automation risk table. - Propose a method to translate rubric-weighted success probabilities into labor-displacement or productivity-gain estimates for a specific industry.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper provides systematic empirical evaluation across 97 market-derived tasks, multiple models, repeated runs, and bootstrap CIs, which yields credible descriptive evidence about current agent capabilities; however, the evaluation relies heavily on an LLM-based judge, a single harness and tool configuration, and a purposive sample of startup-derived workflows, which limit external validity and introduce potential biases. Methods Rigormedium — The authors implement a detailed, reproducible pipeline (startup sampling, user interviews, expert task construction, multi-stage QC, pilot calibration), evaluate multiple models with repeated runs and CIs, and use fine-grained rubrics; but primary evaluation uses an LLM judge (GPT-5.5) which can introduce circularity/bias, task selection uses funding as a coarse filter (potential selection bias), and the evaluation is constrained to one harness, toolset, and a 200-step cap. SampleBenchmark of 97 end-to-end workflow tasks derived from market-validated AI-native startups (20+ startup agents surveyed), grounded by 30+ interviews with deep users and constructed/validated by 50+ domain experts; tasks span 6 domains (Medical & Healthcare, Finance, Legal, Business & Management, STEM & CS, Education & Humanities) and multiple output formats (DOCX, XLSX, PPTX, PDF, Markdown, images, text). Evaluated 9 representative models (closed- and open-source) under a unified Nanobot agent harness with 3 runs per model; rubric-level evaluation performed by AgentJudge (GPT-5.5) and aggregated into task scores and success rates. Themesproductivity human_ai_collab adoption innovation GeneralizabilityTasks are selected from AI-native startups with >$1M funding and evidence of adoption, so workflows may not represent the full distribution of real-world organizational tasks (selection bias)., Evaluation uses a single agent harness (Nanobot) and fixed tool configuration; results may change with different orchestration, plugins, or human-in-the-loop setups., Automatic judging uses an LLM (GPT-5.5) as the primary evaluator, which can bias assessments and may not match human professional judgments in some domains., The 200-interaction-step cap and other operational constraints may systematically disadvantage some long-horizon or iterative workflows., Domain coverage, while broad, is uneven (e.g., only 7 tasks in Education & Humanities), limiting domain-specific generalizability.

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The strongest evaluated general-purpose agents successfully complete only about 30% of StartupBench tasks under the benchmark's strict acceptance criterion. Task Completion Time negative End-to-end task completion rate
Reading fidelity high
Study strength medium
n=97
approximately 30% task completion
0.18
Kimi-K3 achieved the highest average StartupBench score among the evaluated models, with a score of 73.67 out of 100. Output Quality positive Importance-weighted mean per-task benchmark score
Reading fidelity high
Study strength medium
n=97
73.67 points on a 0–100 scale
0.18
GPT-5.6-sol had the highest reported success rate, completing 31.27% of pooled runs at the score threshold of 90 or higher. Output Quality positive Task success rate
Reading fidelity high
Study strength medium
n=291
31.27% success rate
0.18
Current agents generally make substantial partial progress on realistic professional workflows even when they fail to meet the benchmark's completion threshold. Output Quality mixed Average task score versus strict end-to-end completion rate
Reading fidelity high
Study strength medium
n=97
average scores roughly 55–75; completion rates below one third
0.18
Performance varies substantially across domains, with Finance, STEM & Computer Science, and Education & Humanities identified as particularly challenging. Output Quality negative Domain-specific benchmark performance
Reading fidelity high
Study strength medium
n=97
0.18
The paper attributes major sources of agent failure primarily to complex instruction following and domain-specific expertise limitations. Error Rate negative Failure causes in end-to-end workflow execution
Reading fidelity high
Study strength low
n=97
0.09
StartupBench consists of 97 real-world workflow tasks spanning 6 top-level professional domains and requiring diverse deliverable formats. Other null_result Benchmark task and domain coverage
Reading fidelity high
Study strength medium
n=97
97 tasks across 6 domains
0.18
Each StartupBench task is evaluated using an average of 25.3 fine-grained rubrics covering 6 dimensions and 3 importance levels. Output Quality positive Granularity of deliverable evaluation
Reading fidelity high
Study strength medium
n=97
25.3 rubrics per task on average
0.18
StartupBench tasks were derived from market-validated AI startup workflows rather than solely from researcher-defined task assumptions. Adoption Rate positive Real-world demand grounding of benchmark tasks
Reading fidelity high
Study strength medium
n=30
over 30 users interviewed; over 50 experts recruited; 20+ startup agents surveyed
0.18

Notes