0 cumulative citations
View corpus contextState-of-the-art AI agents make substantial partial progress but complete fewer than one-third of real, market-validated workflows; shortcomings in complex instruction-following and domain expertise keep many professional deliverables out of reliable reach.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.
Summary
Main Finding
StartupBench—an end-to-end benchmark built from market-validated AI startup workflows—shows that current general-purpose agent models make substantial partial progress on realistic professional tasks but reliably complete only a minority. Even the strongest evaluated models achieve average rubric-weighted scores in the low-to-mid 70s (out of 100) but succeed (score ≥ 90) on only ~30% or fewer tasks. Key failure modes include complex instruction following and gaps in domain-specific expertise, indicating that many commercially relevant workflows remain beyond reliable automation by generalist agents.
Project page: https://startupbench.github.io/ (paper dated Aug 19, 2026)
Key Points
- Task sourcing: Tasks are derived from workflows in AI-native startups with demonstrated adoption (funding > USD 1M and paid/traction evidence), supplemented by 30+ user interviews and domain-expert reconstruction.
- Benchmark scale & scope:
- 97 end-to-end workflow tasks across 6 domains: Medical & Healthcare (21), Finance (18), Legal (16), Business & Management (19), STEM & CS (16), Education & Humanities (7).
- Multi-format deliverables (DOCX, XLSX, PPTX, PDF, Markdown, images, code, etc.).
- Rich evaluation: average 25.3 fine-grained rubric items per task, covering functionality, structure, formatting, and domain quality.
- Task format: each task is a triple (q: user request, E: workspace/files, R: weighted rubric items). Evaluation aggregates binary rubric outcomes into a weighted score.
- Evaluation methodology:
- Agent-as-a-Judge: each rubric item judged by a lightweight judge agent (GPT-5.5 used as underlying judge) that inspects deliverables and evidence views.
- Models run under a common agent harness (Nanobot), capped at 200 interaction steps.
- 3 independent runs per model; 95% bootstrap CIs reported.
- Models benchmarked: 9 representative models (closed- and open-source families), e.g., Kimi-K3, GPT-5.6-sol, GPT-5.5, Seed-2.1-Pro, Qwen-3.6-Max, DeepSeek-V4-Pro, GLM-5.1, etc.
- Performance summary:
- Best average scores: Kimi-K3 73.67, GPT-5.6-sol 73.61, GPT-5.5 72.79.
- Success rate (score ≥ 90): top models ~26–31%; many domains (Finance, STEM, Education) are especially challenging.
- Failure analysis: primary sources of failure were complex instruction following, nuanced domain expertise, and domain-specific constraints/formatting — not just basic language understanding or single-step reasoning.
Data & Methods
- Data construction pipeline:
- Startup survey to collect candidate agent products (20+ startups as seed pool).
- User interviews (30+ deep users) to capture real workflows, inputs, outputs, constraints.
- Domain experts (50+) converted interview-derived workflows into reproducible tasks and instances; tasks retained only if realistic, end-to-end, and evaluable.
- Multi-stage quality control: expert cross-validation, pilot runs on frontier models (to calibrate difficulty and ensure discriminative value), rubric verification, and manual review.
- Task filtering: tasks that were trivially solved by pilot models were removed to preserve discriminative power.
- Evaluation harness:
- Nanobot agent harness with uniform tool configuration; 200-step interaction cap per task.
- Deliverable evidence view V(D): extracted text + rendered pages to aid judge agents.
- AgentJudge per rubric: returns binary pass/fail and textual justification; final score is weighted aggregation of rubric outcomes.
- Reliability measures: three independent runs per model and 10,000-sample bootstrap for 95% CIs.
- Limitations noted by authors:
- Startup selection bias (funding threshold → favors funded startups).
- Judge agent uses GPT-5.5 which may introduce evaluator model bias.
- Tasks were revised to be reproducible/evaluable which may prune some highly idiosyncratic real-world cases.
- Evaluation limited to a unified agent harness and general-purpose agents; specialized vertical models/agents might perform differently.
Implications for AI Economics
- Grounding benchmarks in market-validated products improves external validity for economic analysis:
- StartupBench measures progress on tasks that users actually pay for, making it more informative for forecasts of automation, productization, and ROI.
- Quantifying automation potential:
- The gap between average rubric scores (~55–75 range) and low success rates (≤ ~30%) suggests substantial partial automation potential but limited full task automation today.
- Economists can use the per-domain success rates to estimate asymmetric automation risk across sectors (e.g., Business & Management easier than Finance/STEM/Education).
- Signals for investors and product strategy:
- The benchmark highlights which professional workflows remain fragile for general-purpose agents—opportunity areas for vertical/specialized agents, domain-tuned models, and toolchain-integrated products.
- Startups that tightly integrate domain expertise, deterministic tooling, and human-in-the-loop flows may capture higher-value use cases where generalists fail.
- Labor market and productivity implications:
- High partial-progress but low completion implies increased complementarities: agents can assist humans (e.g., drafts, analyses) but human oversight remains essential for final deliverables—supporting augmentation rather than immediate displacement in many roles.
- Sectoral heterogeneity in task completion suggests uneven labor impacts; high-skill domain expertise tasks (Finance, STEM) are less automatable now, shifting near-term productivity gains toward workflows amenable to partial automation.
- Valuation and adoption risk:
- Benchmarked evidence may help investors/governance evaluate vendor claims; empirical, rubric-level performance is a useful signal to reduce technological overclaim risk.
- Procurement decisions for enterprises can be better informed by end-to-end success rates and rubric-level failure modes (e.g., instruction following vs. factual knowledge).
- Policy & regulation:
- Demonstrated limitations in domain expertise and regulatory-constrained workflows (e.g., healthcare, legal) argue for caution in delegating critical decisions to generalist agents and for designing certification/validation regimes.
- Research & market implications:
- Indicates economic value in investing in domain-specialized models, structured tool access, better instruction-following, and evaluation pipelines that reflect market needs.
- StartupBench itself can serve as an ongoing empirical indicator: improvements in mean scores and success rates over time can be mapped to likely pace of product-level automation and economic displacement/augmentation.
Potential uses for AI economists: - Map StartupBench success rates to estimates of task-level automation probabilities and sectoral exposure. - Use rubric-level failure modes to forecast which complementary investments (human training, verification tooling, domain data) yield highest productivity returns. - Evaluate startup technology claims and guide investment due diligence using E2E, market-grounded performance rather than isolated capability metrics.
If you want, I can: - Convert the benchmark results into a simple sector-by-sector automation risk table. - Propose a method to translate rubric-weighted success probabilities into labor-displacement or productivity-gain estimates for a specific industry.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The strongest evaluated general-purpose agents successfully complete only about 30% of StartupBench tasks under the benchmark's strict acceptance criterion. Task Completion Time | negative | End-to-end task completion rate |
Reading fidelity
high
Study strength
medium
|
n=97
approximately 30% task completion
|
| Kimi-K3 achieved the highest average StartupBench score among the evaluated models, with a score of 73.67 out of 100. Output Quality | positive | Importance-weighted mean per-task benchmark score |
Reading fidelity
high
Study strength
medium
|
n=97
73.67 points on a 0–100 scale
|
| GPT-5.6-sol had the highest reported success rate, completing 31.27% of pooled runs at the score threshold of 90 or higher. Output Quality | positive | Task success rate |
Reading fidelity
high
Study strength
medium
|
n=291
31.27% success rate
|
| Current agents generally make substantial partial progress on realistic professional workflows even when they fail to meet the benchmark's completion threshold. Output Quality | mixed | Average task score versus strict end-to-end completion rate |
Reading fidelity
high
Study strength
medium
|
n=97
average scores roughly 55–75; completion rates below one third
|
| Performance varies substantially across domains, with Finance, STEM & Computer Science, and Education & Humanities identified as particularly challenging. Output Quality | negative | Domain-specific benchmark performance |
Reading fidelity
high
Study strength
medium
|
n=97
|
| The paper attributes major sources of agent failure primarily to complex instruction following and domain-specific expertise limitations. Error Rate | negative | Failure causes in end-to-end workflow execution |
Reading fidelity
high
Study strength
low
|
n=97
|
| StartupBench consists of 97 real-world workflow tasks spanning 6 top-level professional domains and requiring diverse deliverable formats. Other | null_result | Benchmark task and domain coverage |
Reading fidelity
high
Study strength
medium
|
n=97
97 tasks across 6 domains
|
| Each StartupBench task is evaluated using an average of 25.3 fine-grained rubrics covering 6 dimensions and 3 importance levels. Output Quality | positive | Granularity of deliverable evaluation |
Reading fidelity
high
Study strength
medium
|
n=97
25.3 rubrics per task on average
|
| StartupBench tasks were derived from market-validated AI startup workflows rather than solely from researcher-defined task assumptions. Adoption Rate | positive | Real-world demand grounding of benchmark tasks |
Reading fidelity
high
Study strength
medium
|
n=30
over 30 users interviewed; over 50 experts recruited; 20+ startup agents surveyed
|