Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· data-analysis

DSBench

UT Dallas / Tencent AI Lab · 2024

540 realistic data-analysis and data-modeling tasks sourced from ModelOff and Kaggle competitions.

GitHub stars
Task type
agentic
Modality
multimodal
Access
open
Size
540 items
License
Metrics
accuracy, Relative Performance Gap

No sample rows available for this dataset.

Open it on Hugging Face ↗