METR · 2024
Seven open-ended ML research-engineering environments comparing agents against 61 human experts.
No sample rows available for this dataset.