Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· research-code

RExBench

Boston University · 2025

Twelve tasks where agents autonomously implement novel research extensions to existing papers' codebases.

GitHub stars
Task type
agentic
Modality
code
Access
gated
Size
12 items
License
MIT
Metrics
success rate, execution success rate, file recall

Samples are not shown — this dataset is gated.

Open it on Hugging Face ↗