Scientific LLM Benchmarks
GitHub

A curated, accuracy-first index

Benchmarks for evaluating LLMs on scientific reasoning & discovery

A curated, searchable collection of benchmarks across the natural sciences and agentic research. Each entry includes its paper, code, dataset, and sample items, checked against the original source.

General28Math25Physics & Astro19Chemistry13Materials10Biology21Agentic30
146
benchmarks
7
domains
91
with datasets
125
with code
146 of 146
AAAR-1.02024
Agentic · research-assistance

Expert-annotated tasks assessing LLMs on equation inference, experiment design, and paper-weakness identification.

Penn State et al.open-ended2,142openHF
ABench-Physics2025
Physics & Astro · graduate

500 hard static plus dynamic-variant physics problems testing reasoning and generalization robustness.

Physics
Zhejiang University / Ant Groupopen-ended500open
AGIEval2023
General · human-exam

Human-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions.

MicrosoftMCQ8,062open
ALDbench2024
Materials · synthesis-QA

Expert open-ended QA benchmark evaluating LLMs on atomic-layer-deposition synthesis for accuracy and specificity.

Materials
Argonne National Laboratoryopen-ended70open
ARB2023
General · multi-domain-reasoning

Advanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric.

BiologyChemistryPhysicsMathematics
DuckAI / Georgia Techopen-ended1,207open
ARC (AI2 Reasoning Challenge)2018
General · science-exam-QA

7,787 grade-school science exam multiple-choice questions split into Easy and Challenge sets.

Allen AI (AI2)MCQ7,787openHF
AstaBench2025
Agentic · research-agent-suite

2,400+ problems across eleven benchmarks evaluating agents over the full scientific research pipeline.

Allen AI (AI2)agentic2,400gatedHF
Astro-QA2025
Physics & Astro · astronomy

About 2,700 bilingual astronomy questions across six types spanning astrophysics, celestial mechanics, and astrometry.

PhysicsAstronomy
ACMIS LabQA2,709open
AstroMLab 12024
Physics & Astro · astronomy

4,425 astronomy multiple-choice questions from Annual Reviews testing LLM astronomical knowledge.

PhysicsAstronomy
AstroMLab CollaborationMCQ3,846openHF
AstroVisBench2025
Physics & Astro · astronomy

Evaluates LLMs on end-to-end astronomy computing workflows and scientific result visualization.

PhysicsAstronomy
NSF-Simons CosmicAI Institutecode-gen432openHF
AtmosSci-Bench2025
Physics & Astro · atmospheric-earth

Atmospheric-science problems spanning dynamics, physics, hydrology, geophysics, and oceanography via templated questions.

PhysicsEarth Science
HKUSTMCQ1,301open
AtomWorld2025
Materials · crystalline-spatial-reasoning

Evaluates LLM spatial reasoning on crystalline materials via ten atomic-structure editing actions across four modeling categories over CIF files with verifiable metrics.

Materials
USTC / Shanghai AI Laboratory / UNSWopen-endedopen
BioCoder2023
Biology · code-generation

Fuzz-tested benchmark evaluating LLMs on generating bioinformatics code with cross-file dependencies and domain knowledge.

Biology
Yale University (Gerstein Lab)code-gen207openHF
BioLLMBench2023
Biology · bioinformatics

2,160 runs evaluating GPT-4, Bard, and LLaMA across 24 bioinformatics tasks and metrics.

Biology
UCLAopen-endedopen
BioMaze2025
Biology · pathway-reasoning

5.1K real-research pathway problems testing LLM reasoning over biological pathways and perturbations.

Biology
Peking UniversityQA5,145openHF
BioML-bench2025
Biology · bioinformatics-agent

AI agents build end-to-end biomedical ML pipelines across protein, omics, imaging, and drug tasks.

BiologyMedicine
ScienceMachineagentic24open
Biomni2025
Biology · bioinformatics-agent

General biomedical agent with Biomni-Eval: 433 instances across 10 reasoning tasks.

BiologyMedicine
Stanfordagentic433openHF
BioMysteryBench2026
Biology · bioinformatics-agent

99 expert-written bioinformatics tasks over raw datasets, judged on the final biological conclusion, not the path.

Biology
Anthropicopen-ended99requestHF
BioPlanner2023
Biology · protocols

BioProt dataset automatically evaluating LLMs on generating biology experimental protocols as executable pseudocode.

Biology
FutureHouse / Francis Crick / Oxfordcode-gen100open
BioProBench2025
Biology · protocols

Tests LLMs on biological-protocol QA, step ordering, error correction, generation, and reasoning across 556K instances.

Biology
Peking UniversityQA5,000openHF
BixBench2025
Biology · bioinformatics-agent

Bioinformatics agents tackle 53 real analysis scenarios with ~300 open-answer research questions.

Biology
FutureHouseagentic205openHF
BLADE2024
Agentic · data-analysis

Twelve datasets and research questions evaluating agents' analytical decisions against expert ground truth.

University of Washingtonagentic12open
C-Eval2023
General · exam

13,948 Chinese multiple-choice questions across 52 disciplines and four difficulty levels.

SJTU / HKUSTMCQ13,948openHF
CellVoyager2025
Biology · bioinformatics-agent

CellBench: 76 scRNA-seq studies test agents predicting which analyses the authors performed.

Biology
Stanfordagenticopen
CHAMP2024
Math · competition

270 high-school competition problems annotated with concepts and problem-specific hints.

Mathematics
MITopen-ended270open
ChemBench2024
Chemistry · knowledge

2,700+ curated questions probing LLM chemical knowledge and reasoning against expert chemists.

Chemistry
LAMALab, University of Jena (Jablonka group)MCQ2,788openHF
ChemCoTBench2025
Chemistry · reasoning

Step-wise chemical reasoning across molecular understanding, editing, optimization, and reaction prediction.

Chemistry
IDEA Research / PKU / CUHKopen-ended1,445gatedHF
ChemEval2024
Chemistry · knowledge

Multi-level chemistry benchmark spanning 42 tasks across four progressive difficulty levels for LLMs.

Chemistry
USTCQA5,120openHF
ChemLLMBench2023
Chemistry · molecular

Eight chemistry tasks including reaction prediction, retrosynthesis, and molecule captioning for LLMs.

Chemistry
University of Notre Dame et al.generation800open
ChemSafetyBench2024
Chemistry · safety

Safety benchmark testing LLM handling of hazardous-chemical property, legality, and synthesis queries.

Chemistry
Peking Universitygeneration30,000
CMPhysBench2025
Physics & Astro · condensed-matter

520+ graduate condensed-matter physics problems with a partial-credit metric; top models score below 30%.

Physics
Chinese Academy of Sciences (IOP)open-ended520openHF
CombiBench2025
Math · formal-proof

100 competition combinatorics problems formalized in Lean4, a domain underrepresented by existing formal benchmarks.

Mathematics
Moonshot AI / Numinaproof100openHF
CORE-Bench2024
Agentic · paper-reproduction

270 tasks from 90 papers testing agents on computationally reproducing published scientific results.

Princetonagentic270openHF
CritPt2025
Physics & Astro · research-level-physics

71 composite research-project challenges and 190 modular checkpoints authored by 50+ active physicists, auto-graded via numerical, symbolic, and code evaluations.

Physics
Argonne National Laboratory / UIUC (50+ physicists, 30+ institutions)open-ended71openHF
Curie2025
Agentic · autonomous-discovery

Framework and 46-question benchmark for rigorous, automated scientific experimentation across four CS domains.

University of Michiganagentic46open
DiscoverPhysics2026
Physics & Astro · physics-discovery

22 simulated worlds with non-standard physics where an agent proposes initial conditions, observes noisy N-body trajectories, and submits a natural-language law plus a Python implementation.

Physics
Princeton / NYU / Flatiron Institute / Polymathic AIagentic22openHF
DiscoveryBench2024
Agentic · autonomous-discovery

264 real plus 903 synthetic tasks for data-driven hypothesis discovery across six domains.

Allen AI (AI2)open-ended264openHF
DiscoveryWorld2024
Agentic · autonomous-discovery

Simulated environment with 120 tasks requiring full cycles of hypothesis, experiment, and analysis.

Allen AI (AI2)agentic120open
DSBench2024
Agentic · data-analysis

540 realistic data-analysis and data-modeling tasks sourced from ModelOff and Kaggle competitions.

UT Dallas / Tencent AI Labagentic540openHF
EMMA2025
General · multimodal-STEM

2,788 problems demanding cross-modal reasoning across mathematics, physics, chemistry, and coding for multimodal models.

ChemistryPhysicsMathematics
CUHK / MicrosoftVQA2,788openHF
EXP-Bench2025
Agentic · experiment-automation

Benchmarks agents on designing, implementing, and analyzing end-to-end AI research experiments from publications.

University of Michiganagentic461openHF
FIMO2023
Math · formal-proof

149 IMO-shortlist problems formalized in Lean with informal statements for olympiad-level theorem proving.

Mathematics
Peking University / Huaweiproof149open
FormalMATH2025
Math · formal-proof

5,560 formally verified Lean4 statements from olympiad to undergraduate across algebra, calculus, and number theory.

Mathematics
SphereLab / M-A-Pproof5,560openHF
FrontierMath2024
Math · frontier

Hundreds of unpublished expert-crafted research-level math problems resistant to guessing.

Mathematics
Epoch AIopen-ended338request
FrontierScience2026
General · expert-level-science

Expert-level science benchmark with an olympiad track (IPhO/IChO/IBO-level) and a rubric-graded PhD-level research track across physics, chemistry, and biology.

BiologyChemistryPhysicsMathematics
OpenAIopen-ended160openHF
GAOKAO-Bench2023
General · exam

2,811 Chinese college-entrance-exam questions across subjects for LLM evaluation.

Fudan UniversityMCQ2,811open
GeneBench-Pro2026
Biology · computational-biology

129 messy, judgment-intensive computational-biology problems across ten genomics domains on realistic datasets.

Biology
OpenAIagentic10requestHF
GeneGPT2023
Biology · genomics

Teaches LLMs to call NCBI Web APIs; adds GeneHop, evaluated on GeneTuring.

Biology
NCBIagentic649open
GeneTuring2023
Biology · genomics

16 genomics tasks, 1,600 questions probing LLM genomic knowledge and reasoning.

Biology
Columbia UniversityQA1,600open
Genome-Bench2025
Biology · genomics-reasoning

3,000+ genome-engineering multiple-choice questions mined from expert forum discussions for evaluating LLM reasoning.

Biology
Princeton / StanfordMCQ661openHF
GenoTEX2024
Biology · genomics

Expert-curated gene-expression tasks: dataset selection, preprocessing, and gene-trait statistical analysis.

Biology
UIUCagentic1,384openHF
GPQA2023
General · graduate-QA

448 expert-written graduate-level biology, physics, and chemistry Google-proof multiple-choice questions.

BiologyChemistryPhysics
NYU / Anthropic / CohereMCQ448gatedHF
Gravity-Bench-v12025
Physics & Astro · physics-discovery

Agents plan budgeted observations of simulated binary-star systems to discover the underlying gravitational physics.

PhysicsAstronomy
University of Torontoagentic206openHF
GSM8K2021
Math · grade-school

8.5K grade-school arithmetic word problems requiring multi-step reasoning.

Mathematics
OpenAIopen-ended1,319openHF
HARDMath2024
Math · applied

Graduate applied-math problems requiring asymptotic and approximation methods where leading LLMs score below 45%.

Mathematics
Harvardopen-ended406open
HARP2024
Math · competition

US competition math problems (AMC/AIME/USAMO) with human-written ground-truth solutions.

Mathematics
UCLopen-ended4,780open
HealthBench2025
Biology · clinical-medical

5,000 multi-turn physician-crafted health conversations scored by rubric criteria for LLM performance and safety.

BiologyMedicinePhysics
OpenAIopen-ended5,000openHF
HeurekaBench2026
Agentic · ai-co-scientist

Builds benchmarks of open-ended research questions grounded in real studies and their code to evaluate end-to-end AI co-scientist agents (instantiated as sc-HeurekaBench in single-cell biology).

Biology
EPFL (MLBio Lab)agenticopen
HiPhO2025
Physics & Astro · olympiad

13 recent physics olympiad exams with step-level grading and human medal-based comparison for (M)LLMs.

PhysicsMathematics
CUHK / Shanghai AI Labopen-ended360openHF
Humanity's Last Exam2025
General · frontier

2,500 expert questions across 100+ subjects at the frontier of human academic knowledge.

CAIS / Scale AIQA2,500gatedHF
IMO-Bench2025
Math · olympiad-proof

Olympiad-level suite testing final answers, proof-writing, grading, and Lean formal proofs, vetted by IMO medalists.

Mathematics
Google DeepMindproof460open
InnovatorBench2025
Agentic · ml-research

Evaluates agents on end-to-end innovative LLM-research tasks like data construction, loss design, and reward design.

GAIR-NLP (SJTU)agentic20openHF
JEEBench2023
General · exam

515 challenging IIT-JEE Advanced physics, chemistry, and math problem-solving questions.

ChemistryPhysics
IIT Delhiopen-ended515openHF
LAB-Bench2024
Biology · biology-research

2,400+ MCQs across literature, figures, databases, protocols, and DNA/protein sequence tasks.

Biology
FutureHouseMCQ1,967openHF
LABBench22026
Biology · biology-research

Successor to LAB-Bench with ~1,900 biology-research tasks in realistic contexts (literature, figures, protocols, databases), giving a sharp difficulty jump over LAB-Bench.

Biology
FutureHouse (Edison Scientific)MCQ1,900gatedHF
Lean Workbook2024
Math · autoformalization

About 57K formal-informal Lean4 problem pairs auto-formalized from competition math.

Mathematics
Shanghai AI Labproof57,231openHF
LeanDojo2023
Math · formal-proof

98,734 theorems and proofs from Lean mathlib with premise annotations for retrieval-augmented theorem proving.

Mathematics
Caltech / NVIDIAproof98,734open
LifeSciBench2026
Biology · life-science-research

750 expert-authored free-response life-science research tasks across seven biological domains, rubric-graded.

Biology
OpenAIopen-ended750
LiveIdeaBench2024
Agentic · idea-generation

Benchmarks scientific idea generation from minimal keyword context across five dimensions: originality, feasibility, fluency, flexibility, and clarity.

Renmin University of ChinagenerationopenHF
LLM-SRBench2025
Physics & Astro · equation-discovery

239 problems across four science domains testing genuine LLM equation discovery over memorized formulas.

Physics
Virginia Tech / CMUgeneration239gatedHF
LLM4Mat-Bench2024
Materials · property-prediction

Largest benchmark for LLM materials property prediction over 1.9M crystals and 45 properties.

Materials
Princeton (Vertaix)QA1,900,000open
LMR-Bench2025
Agentic · paper-reproduction

28 tasks from 23 NLP papers testing whether agents reproduce masked research code, verified by unit tests.

UT Dallascode-gen28open
MaCBench2024
Chemistry · multimodal

1,100+ image-question pairs probing multimodal LLM limitations across chemistry and materials workflows.

ChemistryMaterials
LAMALab (Jena) / IIT DelhiVQA1,153openHF
MaScQA2024
Materials · knowledge

650 undergraduate-level materials science questions across 14 categories for evaluating LLM knowledge.

Materials
IIT Delhi (M3RG)MCQ650open
MatCha2025
Materials · multimodal-characterization

1,500 expert questions across 21 tasks for materials characterization image understanding by multimodal models.

Materials
CUHK-ShenzhenVQA1,500openHF
MATH2021
Math · competition

12,500 competition mathematics problems (AMC/AIME level) with step-by-step solutions.

Mathematics
UC Berkeleyopen-ended5,000openHF
MATH-5002023
Math · competition

500-problem held-out subset of MATH widely used for LLM evaluation.

Mathematics
OpenAIopen-ended500openHF
MathBench2024
Math · multi-level

Hierarchical bilingual benchmark spanning arithmetic to college-level theory and application.

Mathematics
Shanghai AI LaboratoryMCQ3,709open
MATHCHECK2024
Math · robustness

Checklist benchmark testing task generalization and reasoning robustness beyond end-to-end answer accuracy.

Mathematics
XJTLU / HKUSTQA4,536openHF
MathOdyssey2024
Math · mixed-level

387 expert-crafted high-school, university, and olympiad-level math problems.

Mathematics
NetMind.AIopen-ended387openHF
MathQA2019
Math · word-problems

37K multiple-choice math word problems annotated with executable operation programs.

Mathematics
UW / Allen AIMCQ2,985openHF
MathVerse2024
Math · multimodal

2,612 visual math problems in six diagram/text variants probing whether MLLMs truly interpret figures.

Mathematics
CUHK MMLab / Shanghai AI LabVQA3,940openHF
MathVista2023
Math · multimodal

Mathematical reasoning benchmark combining visual contexts (figures, charts, diagrams) with problems.

Mathematics
UCLA / University of Washington / MicrosoftVQA6,141openHF
MatSci-NLP2023
Materials · knowledge

Seven materials-science NLP tasks evaluating language models via a unified text-to-schema approach.

Materials
Mila / Universite de Montrealgenerationopen
MatSciBench2025
Materials · reasoning

College-level materials-science reasoning benchmark of 1,340 problems spanning six fields, including multimodal questions.

Materials
UCLAopen-ended1,340openHF
MatText2024
Materials · property-prediction

Benchmarking framework evaluating language models on crystal property prediction across nine text representations.

Materials
LamaLab (Jena) / Intel LabsQAopenHF
MatViX2024
Materials · info-extraction

Benchmarks vision-language models on multimodal information extraction from visually-rich materials-science articles.

Materials
Duke Universitygeneration324open
MatVQA2025
Materials · multimodal-VQA

1,325 questions testing multimodal models on materials imagery like microscopy and diffraction with multi-step reasoning.

Materials
Mila / Universite de MontrealVQA1,325open
miniF2F2021
Math · formal-proof

488 olympiad-level (AMC/AIME/IMO) problems formalized across multiple proof assistants.

Mathematics
OpenAIproof488openHF
MLAgentBench2023
Agentic · ml-research

13 ML experimentation tasks where agents read, write, and execute code to improve performance.

Stanfordagentic13open
MLE-bench2024
Agentic · ml-research

75 Kaggle ML-engineering competitions testing agents against human leaderboards.

OpenAIagentic75open
MLGym2025
Agentic · ml-research

Gym framework with 13 open-ended AI research tasks spanning vision, NLP, RL, and game theory.

Metaagentic13open
MLR-Bench2025
Agentic · ml-research

201 workshop-derived tasks evaluating agents across idea generation, experimentation, and paper writing.

National University of Singaporeagentic201openHF
MLRC-Bench2025
Agentic · ml-research

Benchmarks agents on proposing and coding novel methods for seven recent ML research competition problems.

University of Michiganagentic7open
MMLU-Pro2024
General · knowledge

12K reasoning-focused ten-option questions across 14 academic and STEM domains.

University of Waterloo (TIGER-Lab)MCQ12,102openHF
MMMU2023
General · multimodal

11.5K college-level multimodal questions across six disciplines and 30 subjects.

IN.AI / University of Waterloo / OSUVQA11,550openHF
MolCap-Arena2024
Chemistry · molecule-captioning

Battle-based benchmark scoring 20+ LLMs on molecule captions augmenting molecular property prediction.

Chemistry
Genentech / UIUCgeneration9,957open
MoleculeQA2024
Chemistry · molecular

62K QA pairs over 23K molecules evaluating factual accuracy of LLM molecular comprehension.

Chemistry
IDEA ResearchMCQ62,000openHF
MolTextQA2024
Chemistry · molecule-text-QA

500K QA pairs over 240K PubChem molecules evaluating molecule-structure-to-text understanding and retrieval.

Chemistry
UIUCMCQ12,092openHF
MOOSE-Chem2024
Chemistry · hypothesis-discovery

Tests whether LLMs rediscover unseen chemistry hypotheses from 51 Nature/Science-level papers given background information.

Chemistry
NTU Singapore / Shanghai AI Labgeneration51open
NewtonBench2025
Physics & Astro · law-discovery

324 tasks across twelve physics domains testing LLM agents discovering scientific laws via interactive experimentation.

Physics
HKUSTagentic324open
OlympiadBench2024
General · olympiad

8,476 olympiad-level bilingual multimodal math and physics problems with expert step annotations.

PhysicsMathematics
Tsinghua University / OpenBMBopen-ended8,476openHF
OlympicArena2024
General · olympiad

11,163 olympiad-level problems across seven disciplines for multi-discipline cognitive reasoning.

Mathematics
SJTU (GAIR)QA11,163openHF
Omni-MATH2024
Math · olympiad

4,428 olympiad-level competition problems across 33+ subdomains and difficulty tiers.

Mathematics
Peking Universityopen-ended4,428openHF
OpenBookQA2018
General · science-exam-QA

Elementary science multiple-choice questions requiring core facts plus broad common-sense knowledge.

Allen AI (AI2)MCQ500openHF
Paper2Code2025
Agentic · paper-reproduction

Multi-agent framework and benchmark generating functional code repositories that reproduce ML research papers.

KAIST / DeepAuto.aicode-gen90openHF
PaperBench2025
Agentic · paper-reproduction

Agents replicate 20 ICML 2024 papers from scratch, graded on 8,316 rubric criteria.

OpenAIagentic20open
PaperMind2026
General · scientific-paper-reasoning

Multimodal benchmark evaluating agent-oriented reasoning and critique over real research papers across seven domains via grounding, experimental interpretation, cross-source evidence, and critical-assessment tasks.

University of Illinois Urbana-ChampaignQAopen
PHYBench2025
Physics & Astro · olympiad

500 original physics problems from high-school to olympiad level with expression-edit-distance scoring.

PhysicsMathematics
Peking Universityopen-ended500openHF
PhysGym2025
Physics & Astro · physics-discovery

97 simulated interactive physics-discovery problems with four controlled levels of prior knowledge.

Physics
KAUSTagentic97openHF
PHYSICS2025
Physics & Astro · graduate

1,297 expert-annotated university-level physics problems across six core areas with automated evaluation.

Physics
Yale Universityopen-ended1,297open
PhysReason2025
Physics & Astro · reasoning

1,200 physics problems with step-level automatic scoring for multi-step physics reasoning.

Physics
Xi'an Jiaotong Universityopen-ended1,200openHF
PRL-Bench2026
Physics & Astro · frontier-physics-research

End-to-end physics research benchmark built from ~100 recent Physical Review Letters papers across five subfields, each turned into a long-horizon task scored by an LLM-as-judge.

Physics
Shanghai Jiao Tong Universityagentic100openHF
ProofNet2023
Math · formal-proof

371 undergraduate theorems for autoformalization and formal proving in Lean.

Mathematics
Yale Universityproof371openHF
ProteinLMBench2024
Biology · proteins

944 verified MCQs assessing LLM comprehension of protein sequences and descriptions.

Biology
Shanghai Jiao Tong UniversityMCQ944openHF
Putnam-AXIOM2025
Math · competition

Putnam competition problems plus programmatically perturbed variations giving contamination-resistant unseen instances.

Mathematics
Stanfordopen-ended522openHF
PutnamBench2024
Math · formal-proof

1,600+ Putnam competition problems formalized in Lean, Isabelle, and Coq.

Mathematics
UT Austinproof1,724openHF
QASC2020
General · science-QA

9,980 grade-school science questions requiring retrieval and composition of two facts.

Allen AI (AI2)MCQ926openHF
RE-Bench2024
Agentic · ml-research

Seven open-ended ML research-engineering environments comparing agents against 61 human experts.

METRagentic7open
ResearchBench2025
Agentic · autonomous-discovery

Benchmarks scientific discovery via inspiration retrieval, hypothesis composition, and ranking across twelve disciplines.

Shanghai AI Lab / NTUopen-endedgatedHF
ResearchCodeBench2025
Agentic · scientific-coding

Challenges LLMs to implement novel contributions from recent ML papers by completing TODO code snippets.

Stanford Universitycode-gen212open
RExBench2025
Agentic · research-code

Twelve tasks where agents autonomously implement novel research extensions to existing papers' codebases.

Boston Universityagentic12gatedHF
ScholarChemQA2024
Chemistry · literature-QA

40K abstract-derived chemistry research questions exposing LLM limits in comprehending scholarly literature.

Chemistry
KAUST / Notre DameMCQ1,050open
SciAgentGym2026
Agentic · scientific-tool-use

Agentic science benchmark pairing an environment of 1,780 domain-specific tools across natural-science disciplines with a tiered suite from elementary tool actions to long-horizon workflows.

Fudan NLP Groupagenticopen
SciAssess2024
General · literature

Scientific literature analysis over biology, chemistry, materials, and medicine at three levels.

BiologyChemistryMaterials
DP Technology (deepmodeling)QAopen
SciBench2023
General · exam

692 open-ended college-level chemistry, physics, and math problems requiring multi-step reasoning.

ChemistryPhysics
UCLAopen-ended692openHF
SciCode2024
Agentic · scientific-coding

80 real research coding problems across 16 natural-science subfields, split into 338 subproblems.

UIUC / Princeton / Argonne National Labcode-gen338openHF
ScienceAgentBench2024
Agentic · data-analysis

102 data-driven discovery tasks from 44 peer-reviewed papers, evaluating agents that write Python programs.

Ohio State (OSU-NLP)code-gen102openHF
ScienceQA2022
General · multimodal

~21K multimodal multiple-choice science questions with lecture and chain-of-thought explanations.

UCLA / Allen AI / ASUMCQ4,241openHF
Scientist-Bench2025
Agentic · autonomous-research

Evaluates fully autonomous idea-to-paper research systems across CV, NLP, data mining, and IR against expert papers.

University of Hong Kongagentic28open
SciEval2023
General · knowledge

~18K multi-level questions testing scientific knowledge across chemistry, physics, and biology.

BiologyChemistryPhysics
Fudan University (OpenDFM)MCQ18,000openHF
SciFact2020
General · claim-verification

1,409 expert-written scientific claims verified against research abstracts, with rationales.

Allen AI (AI2)QA1,409openHF
SciKnowEval2024
General · knowledge

Tens of thousands of problems across five cognitive levels in biology, chemistry, physics, and materials.

BiologyChemistryPhysicsMaterials
Zhejiang University (HICAI)QA28,392openHF
SciQ2017
General · science-exam-QA

13,679 crowdsourced science multiple-choice questions across physics, chemistry, and biology with evidence.

BiologyChemistryPhysics
Allen AI (AI2)MCQ1,000openHF
SciReplicate-Bench2025
Agentic · paper-reproduction

Benchmarks agents on reproducing executable code for algorithms described in 36 recent research papers.

King's College Londoncode-gen100open
SDE2025
General · scientific-discovery

Scenario-grounded scientific-discovery benchmark of 43 scenarios and 1,125 questions across biology, chemistry, materials, and physics, plus 8 project-level hypothesis, experiment-design, and interpretation tasks.

BiologyChemistryPhysicsMaterials
Cornell / Princeton / Stanford / MIT / Torontoopen-ended1,125open
SMolInstruct / LlaSMol2024
Chemistry · molecular

Large-scale instruction dataset of 14 small-molecule chemistry tasks used to train and evaluate LlaSMol.

Chemistry
OSU NLP Groupgeneration3,000,000openHF
SUPER2024
Agentic · paper-reproduction

Evaluates agents on setting up and executing tasks from research code repositories.

Allen AI (AI2)agentic799openHF
SuperGPQA2025
General · graduate-QA

26,529 graduate-level questions spanning 285 disciplines, including under-evaluated long-tail fields.

ByteDance Seed / M-A-PMCQ26,529openHF
TheoremQA2023
General · theorem-driven

800 questions applying 350+ theorems across math, physics, EE/CS, and finance.

PhysicsMathematics
University of Waterloo (TIGER-Lab)QA800openHF
TOMG-Bench2024
Chemistry · molecule-generation

Evaluates LLMs on text-based open-domain molecule generation across editing, optimization, and customized generation.

Chemistry
HK PolyU / Shanghai AI Labgeneration45,000openHF
TPBench2025
Physics & Astro · theoretical

57 theoretical physics problems in high-energy theory and cosmology, undergraduate to research level.

PhysicsAstronomy
University of Wisconsin-Madisonopen-ended57requestHF
UGPhysics2025
Physics & Astro · undergraduate

5,520 bilingual undergraduate physics problems across 13 subjects with rule-based judging.

Physics
HKUSTopen-ended5,520openHF
VirBench2026
Biology · agentic-genomics-retrieval

120 curated viral-sequence retrieval queries across ~40 pathogens testing whether LLM agents pull correct ground-truth sequence data from NCBI Virus.

Biology
Anthropic (with NCBI, Broad Institute, Pachter Lab)agentic120request
We-Math2024
Math · multimodal

6.5K visual math problems decomposed into 67 hierarchical knowledge concepts to diagnose reasoning versus memorization.

Mathematics
BUPT / TencentMCQ1,740openHF
Xiezhi2023
General · knowledge

249,587 questions spanning 516 disciplines across 13 categories, continuously updated.

Fudan UniversityMCQ249,587open