Scientific LLM Benchmarks
GitHub
← All benchmarks
Math· frontier

FrontierMath

Epoch AI · 2024

Hundreds of unpublished expert-crafted research-level math problems resistant to guessing.

Mathematics
Task type
open-ended
Modality
text
Access
request
Size
338 items
License
Metrics
accuracy

Examples

Samples are not shown — this dataset is request.