Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· autonomous-research

Scientist-Bench

University of Hong Kong · 2025

Evaluates fully autonomous idea-to-paper research systems across CV, NLP, data mining, and IR against expert papers.

GitHub stars
Task type
agentic
Modality
code
Access
open
Size
28 items
License
Metrics
completeness, correctness (1-5), rating (-3 to 3), comparable %

Examples

No sample rows available for this dataset.