Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· paper-reproduction

Paper2Code

KAIST / DeepAuto.ai · 2025

Multi-agent framework and benchmark generating functional code repositories that reproduce ML research papers.

GitHub stars
Task type
code-gen
Modality
code
Access
open
Size
90 items
License
Apache-2.0
Metrics
correctness (1-5), validity
paper
Generative Judge for Evaluating Alignment
source
poster
repo_name
auto-j
repo_url
https://github.com/GAIR-NLP/auto-j
paper_json
{"abstract":"The rapid development of Large Language Models (LLMs) has substantially expanded the ra
paper_cleaned_json
{"abstract":"The rapid development of Large Language Models (LLMs) has substantially expanded the ra
conference
iclr2024
paper
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
source
poster
repo_name
hidden-context
repo_url
https://github.com/cassidylaidlaw/hidden-context
paper_json
{"abstract":"In practice, preference learning from human feedback depends on incomplete data with hi
paper_cleaned_json
{"abstract":"In practice, preference learning from human feedback depends on incomplete data with hi
conference
iclr2024

Real rows from the Hugging Face datasets server · long values truncated