Hard benchmark for agentic startup work
StartupBench
No frontier agent averages 90; the best pass@1 is only 31.27%.
StartupBench tests whether general-purpose agents can turn messy, long-horizon startup workspaces into production-ready spreadsheets, documents, slides, reports, and other deliverables.
Leaderboard
StartupBench is hard: every model misses the production-ready line.
Scores are averaged over realistic long-horizon workflows. Most model averages sit roughly between 55 and 75, the best average score is 73.67, and the best pass@1 is only 31.27 under the strict score >= 90 criterion.
| Rank | Model | Average | Pass@1 | Gap to pass |
|---|
Top average: 73.67. Top pass@1: 31.27%. No evaluated model reaches the 90-point production-ready line.
Case explorer
Explore concrete cases.
Each case opens into a compact replay view with the task, full trajectory, model scores, submitted result, and failure points behind the leaderboard.
Benchmark anatomy
Six domains, many deliverable formats.
The benchmark stresses professional conventions across medical and healthcare, finance, legal, management, STEM, and education workflows, with outputs ranging from spreadsheets to slide decks and reports.
Project overview
From isolated tests to delegated work.
Market-sourced tasks
Tasks are derived from startup-agent products, demos, and user interviews around workflows people already want to delegate.
End-to-end artifacts
Agents must plan, inspect files, compute, write, format, and submit original deliverables rather than short text answers.
Rubric-level judging
Submissions are evaluated through evidence views and weighted rubrics, with a 92% agreement rate against expert judgments.
rubrics per task
pass threshold
human agreement
average framework variation
Benchmark construction
How the workflow tasks are synthesized.
The synthesized construction pipeline is kept after the main performance, case, and benchmark sections.
Citation
StartupBench paper.
@misc{zhu2026startupbench,
title = {StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows},
author = {Liya Zhu and Xin Ma and Tao Liu and Haodong Wang and Ge Zhang and Jingzhe Ding and Qingshui Gu and Yongjie Zhong and Jinxiang Meng and Yuan Gao and Yunqiu Zhou and Hao Zhu and others},
year = {2026},
eprint = {2608.17800},
archivePrefix = {arXiv},
primaryClass = {cs.AI}
}