Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· ml-research

RE-Bench

METR · 2024

Seven open-ended ML research-engineering environments comparing agents against 61 human experts.

GitHub stars
Task type
agentic
Modality
code
Access
open
Size
7 items
License
MIT
Metrics
normalized score, score@k

Examples

No sample rows available for this dataset.