Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· paper-reproduction

SciReplicate-Bench

King's College London · 2025

Benchmarks agents on reproducing executable code for algorithms described in 36 recent research papers.

GitHub stars
Task type
code-gen
Modality
code
Access
open
Size
100 items
License
Metrics
CodeBLEU, execution accuracy, recall

Examples

No sample rows available for this dataset.