Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· paper-reproduction

CORE-Bench

Princeton · 2024

270 tasks from 90 papers testing agents on computationally reproducing published scientific results.

GitHub stars
Task type
agentic
Modality
code
Access
open
Size
270 items
License
MIT
Metrics
accuracy

No sample rows available for this dataset.

Open it on Hugging Face ↗