Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· ml-research

MLGym

Meta · 2025

Gym framework with 13 open-ended AI research tasks spanning vision, NLP, RL, and game theory.

GitHub stars
Task type
agentic
Modality
code
Access
open
Size
13 items
License
CC-BY-NC-4.0
Metrics
AUP, performance profile

Examples

No sample rows available for this dataset.