◆Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· interactive-science-env

ScienceWorld

University of Arizona / Microsoft Research / Allen AI (AI2) · 2022

Interactive text environment with 30 elementary-science tasks where agents must run the experiment — melt a substance, test conductivity, breed a plant — instead of reciting the answer.

GitHub stars
Task type
agentic
Modality
text
Access
open
Size
30 items
License
Apache-2.0
Metrics
task score

Examples

No sample rows available for this dataset.