Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· paper-reproduction

LMR-Bench

UT Dallas · 2025

28 tasks from 23 NLP papers testing whether agents reproduce masked research code, verified by unit tests.

GitHub stars
Task type
code-gen
Modality
code
Access
open
Size
28 items
License
MIT
Metrics
accuracy, LLM-as-judge

Examples

No sample rows available for this dataset.