UT Dallas · 2025
28 tasks from 23 NLP papers testing whether agents reproduce masked research code, verified by unit tests.
No sample rows available for this dataset.