Stanford · 2023
13 ML experimentation tasks where agents read, write, and execute code to improve performance.
No sample rows available for this dataset.