◆Scientific LLM Benchmarks
GitHub
← All benchmarks
Math· uncontaminated-competition

MathArena

ETH Zurich / INSAIT · 2025

Evaluates models only on competitions held after their release date to rule out contamination, covering 162 problems from seven competitions plus human-graded IMO proof writing.

Mathematics
GitHub stars
Task type
open-ended
Modality
text
Access
open
Size
162 items
License
MIT
Metrics
accuracy

Examples

No sample rows available for this dataset.