Hard benchmark for agentic startup work

StartupBench

No frontier agent averages 90; the best pass@1 is only 31.27%.

StartupBench tests whether general-purpose agents can turn messy, long-horizon startup workspaces into production-ready spreadsheets, documents, slides, reports, and other deliverables.

StartupBench workflow overview from user request to multi-format deliverables.

Leaderboard

StartupBench is hard: every model misses the production-ready line.

Scores are averaged over realistic long-horizon workflows. Most model averages sit roughly between 55 and 75, the best average score is 73.67, and the best pass@1 is only 31.27 under the strict score >= 90 criterion.

Model scores on StartupBench

Average score and pass@1 across evaluated frontier agents.

Rank Model Average Pass@1 Gap to pass

Top average: 73.67. Top pass@1: 31.27%. No evaluated model reaches the 90-point production-ready line.

97 real-world workflow tasks
6 professional domains
25.3 rubrics per task on average
9 frontier agents evaluated

Case explorer

Explore concrete cases.

Each case opens into a compact replay view with the task, full trajectory, model scores, submitted result, and failure points behind the leaderboard.

Benchmark anatomy

Six domains, many deliverable formats.

The benchmark stresses professional conventions across medical and healthcare, finance, legal, management, STEM, and education workflows, with outputs ranging from spreadsheets to slide decks and reports.

Medical & HealthCare21
Finance18
Business & Management19
Legal16
STEM & Computer Science16
Education & Humanities7
Table 2 statistics of StartupBench.
DOCX XLSX PPTX PDF Markdown Images Text artifacts

Project overview

From isolated tests to delegated work.

01

Market-sourced tasks

Tasks are derived from startup-agent products, demos, and user interviews around workflows people already want to delegate.

02

End-to-end artifacts

Agents must plan, inspect files, compute, write, format, and submit original deliverables rather than short text answers.

03

Rubric-level judging

Submissions are evaluated through evidence views and weighted rubrics, with a 92% agreement rate against expert judgments.

25.3

rubrics per task

90

pass threshold

92%

human agreement

1.66

average framework variation

Benchmark construction

How the workflow tasks are synthesized.

The synthesized construction pipeline is kept after the main performance, case, and benchmark sections.

StartupBench construction pipeline from startup-agent selection to quality control and difficulty calibration.

Citation

StartupBench paper.

@misc{zhu2026startupbench,
  title         = {StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows},
  author        = {Liya Zhu and Xin Ma and Tao Liu and Haodong Wang and Ge Zhang and Jingzhe Ding and Qingshui Gu and Yongjie Zhong and Jinxiang Meng and Yuan Gao and Yunqiu Zhou and Hao Zhu and others},
  year          = {2026},
  eprint        = {2608.17800},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI}
}