Scientific LLM Benchmarks
GitHub
← All benchmarks
Chemistry· knowledge

ChemEval

USTC · 2024

Multi-level chemistry benchmark spanning 42 tasks across four progressive difficulty levels for LLMs.

Chemistry
GitHub stars
Task type
QA
Modality
multimodal
Access
open
Size
5,120 items
License
CC-BY-NC-4.0
Metrics
accuracy, LLM-based grading
query
Please convert the given SMILES(CC.CC1=CC=CC=C1)to SELFIES.You must output your prediction, i.e. a valid SELFIES, and follow the output format exactly as follows: {"answer": "The answer you judge "}. I don't need any explanation, you just need to output your judgment in format.
target
[C][C].[C][C][=C][C][=C][C][=C][Ring1][=Branch1]
filename
SMILES与SELFIES互译_test.json
query
Please convert the given SMILES(CCOC(=O)C)to SELFIES.You must output your prediction, i.e. a valid SELFIES, and follow the output format exactly as follows: {"answer": "The answer you judge "}. I don't need any explanation, you just need to output your judgment in format.
target
[C][C][O][C][=Branch1][C][=O][C]
filename
SMILES与SELFIES互译_test.json

Real rows from the Hugging Face datasets server · long values truncated