◆Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· research-agent-suite

AstaBench

Allen AI (AI2) · 2025

ICLR 2026 oral suite with 2,400+ problems across eleven benchmarks covering literature, coding, data analysis, and end-to-end scientific research.

GitHub stars
Task type
agentic
Modality
multimodal
Access
gated
Size
2,400 items
License
Apache-2.0
Metrics
accuracy

Samples are not shown — this dataset is gated.

Open it on Hugging Face ↗