COREDesignerAPI

onepot-Bench 0

A chemistry benchmark suite for language models, built to measure whether a model can make the decisions necessary to execute chemical synthesis experiments.

ChemAbacus consists of 800 QA questions, each concerning a single molecule given as a SMILES string. ChemAbacus measures general chemistry literacy and numeracy: strong performance on its tasks is likely necessary but not sufficient for real-world chemistry decision-making.

View
Refusals
Provider
Accuracy against cost
758188941000.3131030Cost per query (¢)Accuracy (%)Claude Opus 5 (high)Claude Fable 5 (low)

Accuracy per query type. MW within ±1%, TPSA within ±5 Ų, cLogP within ±1; rings, rotatable bonds, HBA and HBD are exact match; FG is Jaccard over the functional-group set.

Cell color
40100%
ModelEffortMWRingsTPSAcLogPRotHBAHBDFGMacro
GPT-5.5low9310090765193989487
GPT-5.5med9899968568941009392
GPT-5.5high99100978570961009392
GPT-Rosalindlow889770695377959681
GPT-Rosalindmed959978777782969688
GPT-Rosalindhigh9910081778089869688
GPT-5.6 Sollow9610080746191649583
GPT-5.6 Solmed9810078786793669684
GPT-5.6 Solhigh9910086737895659686
Claude Opus 4.8low9299967590881009692
Claude Opus 4.8med96100997990921009794
Claude Opus 4.8high981001007893901009794
Claude Opus 5low95100967287901009692
Claude Opus 5med97100998294901009695
Claude Opus 5high10010099879396979596
Claude Fable 5low879191667981969586·38r
Claude Fable 5med918987698382969586·60r
Claude Fable 5high918486788583979688·75r

Nr denotes number of refusals.