Scientific LLM Benchmarks
GitHub
← All benchmarks
General· frontier

Humanity's Last Exam

CAIS / Scale AI · 2025

2,500 expert questions across 100+ subjects at the frontier of human academic knowledge.

GitHub stars
Task type
QA
Modality
multimodal
Access
gated
Size
2,500 items
License
MIT
Metrics
accuracy, calibration error

Samples are not shown — this dataset is gated.

Open it on Hugging Face ↗