Scientific LLM Benchmarks
GitHub
← All benchmarks
General· scientific-paper-reasoning

PaperMind

University of Illinois Urbana-Champaign · 2026

Multimodal benchmark evaluating agent-oriented reasoning and critique over real research papers across seven domains via grounding, experimental interpretation, cross-source evidence, and critical-assessment tasks.

GitHub stars
Task type
QA
Modality
multimodal
Access
open
Size
License
Metrics
accuracy

Examples

No sample rows available for this dataset.