Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· ml-research

MLAgentBench

Stanford · 2023

13 ML experimentation tasks where agents read, write, and execute code to improve performance.

GitHub stars
Task type
agentic
Modality
code
Access
open
Size
13 items
License
MIT
Metrics
success rate, average improvement

Examples

No sample rows available for this dataset.