Expert-annotated tasks assessing LLMs on equation inference, experiment design, and paper-weakness identification.
500 hard static plus dynamic-variant physics problems testing reasoning and generalization robustness.
Human-centric benchmark from standardized exams like Gaokao, SAT, LSAT, and math competitions.
Expert open-ended QA benchmark evaluating LLMs on atomic-layer-deposition synthesis for accuracy and specificity.
Advanced problems across mathematics, physics, biology, chemistry, and law, including symbolic and proof items graded by rubric.
7,787 grade-school science exam multiple-choice questions split into Easy and Challenge sets.
2,400+ problems across eleven benchmarks evaluating agents over the full scientific research pipeline.
About 2,700 bilingual astronomy questions across six types spanning astrophysics, celestial mechanics, and astrometry.
4,425 astronomy multiple-choice questions from Annual Reviews testing LLM astronomical knowledge.
Evaluates LLMs on end-to-end astronomy computing workflows and scientific result visualization.
Atmospheric-science problems spanning dynamics, physics, hydrology, geophysics, and oceanography via templated questions.
Evaluates LLM spatial reasoning on crystalline materials via ten atomic-structure editing actions across four modeling categories over CIF files with verifiable metrics.
Fuzz-tested benchmark evaluating LLMs on generating bioinformatics code with cross-file dependencies and domain knowledge.
2,160 runs evaluating GPT-4, Bard, and LLaMA across 24 bioinformatics tasks and metrics.
5.1K real-research pathway problems testing LLM reasoning over biological pathways and perturbations.
AI agents build end-to-end biomedical ML pipelines across protein, omics, imaging, and drug tasks.
General biomedical agent with Biomni-Eval: 433 instances across 10 reasoning tasks.
99 expert-written bioinformatics tasks over raw datasets, judged on the final biological conclusion, not the path.
BioProt dataset automatically evaluating LLMs on generating biology experimental protocols as executable pseudocode.
Tests LLMs on biological-protocol QA, step ordering, error correction, generation, and reasoning across 556K instances.
Bioinformatics agents tackle 53 real analysis scenarios with ~300 open-answer research questions.
Twelve datasets and research questions evaluating agents' analytical decisions against expert ground truth.
13,948 Chinese multiple-choice questions across 52 disciplines and four difficulty levels.
CellBench: 76 scRNA-seq studies test agents predicting which analyses the authors performed.
270 high-school competition problems annotated with concepts and problem-specific hints.
2,700+ curated questions probing LLM chemical knowledge and reasoning against expert chemists.
Step-wise chemical reasoning across molecular understanding, editing, optimization, and reaction prediction.
Multi-level chemistry benchmark spanning 42 tasks across four progressive difficulty levels for LLMs.
Eight chemistry tasks including reaction prediction, retrosynthesis, and molecule captioning for LLMs.
Safety benchmark testing LLM handling of hazardous-chemical property, legality, and synthesis queries.
520+ graduate condensed-matter physics problems with a partial-credit metric; top models score below 30%.
100 competition combinatorics problems formalized in Lean4, a domain underrepresented by existing formal benchmarks.
270 tasks from 90 papers testing agents on computationally reproducing published scientific results.
71 composite research-project challenges and 190 modular checkpoints authored by 50+ active physicists, auto-graded via numerical, symbolic, and code evaluations.
Framework and 46-question benchmark for rigorous, automated scientific experimentation across four CS domains.
22 simulated worlds with non-standard physics where an agent proposes initial conditions, observes noisy N-body trajectories, and submits a natural-language law plus a Python implementation.
264 real plus 903 synthetic tasks for data-driven hypothesis discovery across six domains.
Simulated environment with 120 tasks requiring full cycles of hypothesis, experiment, and analysis.
540 realistic data-analysis and data-modeling tasks sourced from ModelOff and Kaggle competitions.
2,788 problems demanding cross-modal reasoning across mathematics, physics, chemistry, and coding for multimodal models.
Benchmarks agents on designing, implementing, and analyzing end-to-end AI research experiments from publications.
149 IMO-shortlist problems formalized in Lean with informal statements for olympiad-level theorem proving.
5,560 formally verified Lean4 statements from olympiad to undergraduate across algebra, calculus, and number theory.
Hundreds of unpublished expert-crafted research-level math problems resistant to guessing.
Expert-level science benchmark with an olympiad track (IPhO/IChO/IBO-level) and a rubric-graded PhD-level research track across physics, chemistry, and biology.
2,811 Chinese college-entrance-exam questions across subjects for LLM evaluation.
129 messy, judgment-intensive computational-biology problems across ten genomics domains on realistic datasets.
Teaches LLMs to call NCBI Web APIs; adds GeneHop, evaluated on GeneTuring.
16 genomics tasks, 1,600 questions probing LLM genomic knowledge and reasoning.
3,000+ genome-engineering multiple-choice questions mined from expert forum discussions for evaluating LLM reasoning.
Expert-curated gene-expression tasks: dataset selection, preprocessing, and gene-trait statistical analysis.
448 expert-written graduate-level biology, physics, and chemistry Google-proof multiple-choice questions.
Agents plan budgeted observations of simulated binary-star systems to discover the underlying gravitational physics.
8.5K grade-school arithmetic word problems requiring multi-step reasoning.
Graduate applied-math problems requiring asymptotic and approximation methods where leading LLMs score below 45%.
US competition math problems (AMC/AIME/USAMO) with human-written ground-truth solutions.
5,000 multi-turn physician-crafted health conversations scored by rubric criteria for LLM performance and safety.
Builds benchmarks of open-ended research questions grounded in real studies and their code to evaluate end-to-end AI co-scientist agents (instantiated as sc-HeurekaBench in single-cell biology).
13 recent physics olympiad exams with step-level grading and human medal-based comparison for (M)LLMs.
2,500 expert questions across 100+ subjects at the frontier of human academic knowledge.
Olympiad-level suite testing final answers, proof-writing, grading, and Lean formal proofs, vetted by IMO medalists.
Evaluates agents on end-to-end innovative LLM-research tasks like data construction, loss design, and reward design.
515 challenging IIT-JEE Advanced physics, chemistry, and math problem-solving questions.
2,400+ MCQs across literature, figures, databases, protocols, and DNA/protein sequence tasks.
Successor to LAB-Bench with ~1,900 biology-research tasks in realistic contexts (literature, figures, protocols, databases), giving a sharp difficulty jump over LAB-Bench.
About 57K formal-informal Lean4 problem pairs auto-formalized from competition math.
98,734 theorems and proofs from Lean mathlib with premise annotations for retrieval-augmented theorem proving.
750 expert-authored free-response life-science research tasks across seven biological domains, rubric-graded.
Benchmarks scientific idea generation from minimal keyword context across five dimensions: originality, feasibility, fluency, flexibility, and clarity.
239 problems across four science domains testing genuine LLM equation discovery over memorized formulas.
Largest benchmark for LLM materials property prediction over 1.9M crystals and 45 properties.
28 tasks from 23 NLP papers testing whether agents reproduce masked research code, verified by unit tests.
1,100+ image-question pairs probing multimodal LLM limitations across chemistry and materials workflows.
650 undergraduate-level materials science questions across 14 categories for evaluating LLM knowledge.
1,500 expert questions across 21 tasks for materials characterization image understanding by multimodal models.
12,500 competition mathematics problems (AMC/AIME level) with step-by-step solutions.
500-problem held-out subset of MATH widely used for LLM evaluation.
Hierarchical bilingual benchmark spanning arithmetic to college-level theory and application.
Checklist benchmark testing task generalization and reasoning robustness beyond end-to-end answer accuracy.
387 expert-crafted high-school, university, and olympiad-level math problems.
37K multiple-choice math word problems annotated with executable operation programs.
2,612 visual math problems in six diagram/text variants probing whether MLLMs truly interpret figures.
Mathematical reasoning benchmark combining visual contexts (figures, charts, diagrams) with problems.
Seven materials-science NLP tasks evaluating language models via a unified text-to-schema approach.
College-level materials-science reasoning benchmark of 1,340 problems spanning six fields, including multimodal questions.
Benchmarking framework evaluating language models on crystal property prediction across nine text representations.
Benchmarks vision-language models on multimodal information extraction from visually-rich materials-science articles.
1,325 questions testing multimodal models on materials imagery like microscopy and diffraction with multi-step reasoning.
488 olympiad-level (AMC/AIME/IMO) problems formalized across multiple proof assistants.
13 ML experimentation tasks where agents read, write, and execute code to improve performance.
75 Kaggle ML-engineering competitions testing agents against human leaderboards.
Gym framework with 13 open-ended AI research tasks spanning vision, NLP, RL, and game theory.
201 workshop-derived tasks evaluating agents across idea generation, experimentation, and paper writing.
Benchmarks agents on proposing and coding novel methods for seven recent ML research competition problems.
12K reasoning-focused ten-option questions across 14 academic and STEM domains.
11.5K college-level multimodal questions across six disciplines and 30 subjects.
Battle-based benchmark scoring 20+ LLMs on molecule captions augmenting molecular property prediction.
62K QA pairs over 23K molecules evaluating factual accuracy of LLM molecular comprehension.
500K QA pairs over 240K PubChem molecules evaluating molecule-structure-to-text understanding and retrieval.
Tests whether LLMs rediscover unseen chemistry hypotheses from 51 Nature/Science-level papers given background information.
324 tasks across twelve physics domains testing LLM agents discovering scientific laws via interactive experimentation.
8,476 olympiad-level bilingual multimodal math and physics problems with expert step annotations.
11,163 olympiad-level problems across seven disciplines for multi-discipline cognitive reasoning.
4,428 olympiad-level competition problems across 33+ subdomains and difficulty tiers.
Elementary science multiple-choice questions requiring core facts plus broad common-sense knowledge.
Multi-agent framework and benchmark generating functional code repositories that reproduce ML research papers.
Agents replicate 20 ICML 2024 papers from scratch, graded on 8,316 rubric criteria.
Multimodal benchmark evaluating agent-oriented reasoning and critique over real research papers across seven domains via grounding, experimental interpretation, cross-source evidence, and critical-assessment tasks.
500 original physics problems from high-school to olympiad level with expression-edit-distance scoring.
97 simulated interactive physics-discovery problems with four controlled levels of prior knowledge.
1,297 expert-annotated university-level physics problems across six core areas with automated evaluation.
1,200 physics problems with step-level automatic scoring for multi-step physics reasoning.
End-to-end physics research benchmark built from ~100 recent Physical Review Letters papers across five subfields, each turned into a long-horizon task scored by an LLM-as-judge.
371 undergraduate theorems for autoformalization and formal proving in Lean.
944 verified MCQs assessing LLM comprehension of protein sequences and descriptions.
Putnam competition problems plus programmatically perturbed variations giving contamination-resistant unseen instances.
1,600+ Putnam competition problems formalized in Lean, Isabelle, and Coq.
9,980 grade-school science questions requiring retrieval and composition of two facts.
Seven open-ended ML research-engineering environments comparing agents against 61 human experts.
Benchmarks scientific discovery via inspiration retrieval, hypothesis composition, and ranking across twelve disciplines.
Challenges LLMs to implement novel contributions from recent ML papers by completing TODO code snippets.
Twelve tasks where agents autonomously implement novel research extensions to existing papers' codebases.
40K abstract-derived chemistry research questions exposing LLM limits in comprehending scholarly literature.
Agentic science benchmark pairing an environment of 1,780 domain-specific tools across natural-science disciplines with a tiered suite from elementary tool actions to long-horizon workflows.
Scientific literature analysis over biology, chemistry, materials, and medicine at three levels.
692 open-ended college-level chemistry, physics, and math problems requiring multi-step reasoning.
80 real research coding problems across 16 natural-science subfields, split into 338 subproblems.
102 data-driven discovery tasks from 44 peer-reviewed papers, evaluating agents that write Python programs.
~21K multimodal multiple-choice science questions with lecture and chain-of-thought explanations.
Evaluates fully autonomous idea-to-paper research systems across CV, NLP, data mining, and IR against expert papers.
~18K multi-level questions testing scientific knowledge across chemistry, physics, and biology.
1,409 expert-written scientific claims verified against research abstracts, with rationales.
Tens of thousands of problems across five cognitive levels in biology, chemistry, physics, and materials.
13,679 crowdsourced science multiple-choice questions across physics, chemistry, and biology with evidence.
Benchmarks agents on reproducing executable code for algorithms described in 36 recent research papers.
Scenario-grounded scientific-discovery benchmark of 43 scenarios and 1,125 questions across biology, chemistry, materials, and physics, plus 8 project-level hypothesis, experiment-design, and interpretation tasks.
Large-scale instruction dataset of 14 small-molecule chemistry tasks used to train and evaluate LlaSMol.
Evaluates agents on setting up and executing tasks from research code repositories.
26,529 graduate-level questions spanning 285 disciplines, including under-evaluated long-tail fields.
800 questions applying 350+ theorems across math, physics, EE/CS, and finance.
Evaluates LLMs on text-based open-domain molecule generation across editing, optimization, and customized generation.
57 theoretical physics problems in high-energy theory and cosmology, undergraduate to research level.
5,520 bilingual undergraduate physics problems across 13 subjects with rule-based judging.
120 curated viral-sequence retrieval queries across ~40 pathogens testing whether LLM agents pull correct ground-truth sequence data from NCBI Virus.
6.5K visual math problems decomposed into 67 hierarchical knowledge concepts to diagnose reasoning versus memorization.
249,587 questions spanning 516 disciplines across 13 categories, continuously updated.