Scientific LLM Benchmarks
GitHub
← All benchmarks
General· human-exam

AGIEval

Microsoft · 2023

Human-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions.

GitHub stars
Task type
MCQ
Modality
text
Access
open
Size
8,062 items
License
MIT
Metrics
accuracy

Examples

No sample rows available for this dataset.