Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· autonomous-discovery

DiscoveryWorld

Allen AI (AI2) · 2024

Simulated environment with 120 tasks requiring full cycles of hypothesis, experiment, and analysis.

GitHub stars
Task type
agentic
Modality
multimodal
Access
open
Size
120 items
License
Apache-2.0
Metrics
task completion, process score, knowledge score

Examples

No sample rows available for this dataset.