Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· paper-reproduction

PaperBench

OpenAI · 2025

Agents replicate 20 ICML 2024 papers from scratch, graded on 8,316 rubric criteria.

Task type
agentic
Modality
code
Access
open
Size
20 items
License
MIT
Metrics
replication score

Examples

No sample rows available for this dataset.