Scientific LLM Benchmarks
GitHub
← All benchmarks
Math· robustness

MATHCHECK

XJTLU / HKUST · 2024

Checklist benchmark testing task generalization and reasoning robustness beyond end-to-end answer accuracy.

Mathematics
GitHub stars
Task type
QA
Modality
multimodal
Access
open
Size
4,536 items
License
CC-BY-4.0
Metrics
accuracy

No sample rows available for this dataset.

Open it on Hugging Face ↗