Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· scientific-coding

ResearchCodeBench

Stanford University · 2025

Challenges LLMs to implement novel contributions from recent ML papers by completing TODO code snippets.

GitHub stars
Task type
code-gen
Modality
code
Access
open
Size
212 items
License
CC-BY-SA-4.0
Metrics
Scaled Pass@1, Pass@1

Examples

No sample rows available for this dataset.