Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· ml-research

MLE-bench

OpenAI · 2024

75 Kaggle ML-engineering competitions testing agents against human leaderboards.

GitHub stars
Task type
agentic
Modality
code
Access
open
Size
75 items
License
MIT
Metrics
medal rate

Examples

No sample rows available for this dataset.