Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· ai-co-scientist

HeurekaBench

EPFL (MLBio Lab) · 2026

Builds benchmarks of open-ended research questions grounded in real studies and their code to evaluate end-to-end AI co-scientist agents (instantiated as sc-HeurekaBench in single-cell biology).

Biology
GitHub stars
Task type
agentic
Modality
code
Access
open
Size
License
Metrics
workflow correctness

Examples

No sample rows available for this dataset.