Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· research-assistance

AAAR-1.0

Penn State et al. · 2024

Expert-annotated tasks assessing LLMs on equation inference, experiment design, and paper-weakness identification.

GitHub stars
Task type
open-ended
Modality
multimodal
Access
open
Size
2,142 items
License
MIT
Metrics
F1, ROUGE-L, BERTScore

No sample rows available for this dataset.

Open it on Hugging Face ↗