Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· scientific-tool-use

SciAgentGym

Fudan NLP Group · 2026

Agentic science benchmark pairing an environment of 1,780 domain-specific tools across natural-science disciplines with a tiered suite from elementary tool actions to long-horizon workflows.

GitHub stars
Task type
agentic
Modality
code
Access
open
Size
License
Metrics
success rate

Examples

No sample rows available for this dataset.