Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· research-agent-suite

AstaBench

Allen AI (AI2) · 2025

2,400+ problems across eleven benchmarks evaluating agents over the full scientific research pipeline.

GitHub stars
Task type
agentic
Modality
multimodal
Access
gated
Size
2,400 items
License
Apache-2.0
Metrics
accuracy

Samples are not shown — this dataset is gated.

Open it on Hugging Face ↗