Scientific LLM Benchmarks
GitHub
← All benchmarks
Agentic· ml-research

MLR-Bench

National University of Singapore · 2025

201 workshop-derived tasks evaluating agents across idea generation, experimentation, and paper writing.

GitHub stars
Task type
agentic
Modality
text
Access
open
Size
201 items
License
MIT
Metrics
LLM-as-judge
task_id
0
task_name
iclr2023_bands
task_description
# Backdoor Attacks and Defenses in Machine Learning ## Overview Backdoor attacks aim to cause consistent misclassification of any input by adding a specific pattern called a trigger. Unlike adversarial attacks requiring generating perturbations on the fly to induce misclassification for one single input, backdoor attacks have prompt effects by simply applying a pre-chosen trigger. Recent studies have shown the feasibility of launching backdoor attacks in various domains, such as computer vision (CV), natural language processing (NLP), federated learning (FL), etc. As backdoor attacks are mos …
task_id
1
task_name
iclr2023_dg
task_description
# What do we need for successful domain generalization? ## Workshop Description The real challenge for any machine learning system is to be reliable and robust in any situation, even if it is different compared to training conditions. Existing general purpose approaches to domain generalization (DG) — a problem setting that challenges a model to generalize well to data outside the distribution sampled at training time — have failed to consistently outperform standard empirical risk minimization baselines. In this workshop, we aim to work towards answering a single question: what do we need f …

Real rows from the Hugging Face datasets server · long values truncated