Scientific LLM Benchmarks
GitHub
← All benchmarks
Biology· bioinformatics-agent

BioMysteryBench

Anthropic · 2026

99 expert-written bioinformatics tasks over raw datasets, judged on the final biological conclusion, not the path.

Biology
Task type
open-ended
Modality
text
Access
request
Size
99 items
License
CC-BY-4.0
Metrics
accuracy, solve consistency (k/5)

Samples are not shown — this dataset is request.

Open it on Hugging Face ↗