Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· ml-research

InnovatorBench

GAIR-NLP (SJTU) · 2025

Evaluates agents on end-to-end innovative LLM-research tasks like data construction, loss design, and reward design.

GitHub stars
Task type
agentic
Modality
code
Access
open
Size
20 items
License
Apache-2.0
Metrics
weighted average score

No sample rows available for this dataset.

Open it on Hugging Face ↗