遇见数据集

Orthonym: verified IUPAC names for chemical structures in the wild - evaluation splits, corpora, run records and Source Data

收藏
Zenodo2026-09-29 更新2026-10-01 收录
官方服务:

资源简介:

Data for the manuscript "Orthonym: verified IUPAC names for chemical structures in the wild" (K. Rajan, A. Zielesny, C. Steinbeck). Orthonym is an open, deterministic, rule-based generator of IUPAC names that emits a name only after OPSIN has read the name back to the input structure. This record holds everything needed to check every number in the paper. The single archive Zenodo_Orthonym_data.zip unpacks to one folder: splits/ - the three ChEBI evaluation splits (dev500, dev2000, holdout2000) as JSON with seed, prefilter, checksum and SMILES; holdout2000 also as a SMILES list. corpora/ - the sets as named: QM9 (133,885), all unique ChEBI structures (111,843), 500,000 PubChem and 500,000 ZINC22 molecules, and the 1,000,000-molecule scaffold-representative PubChem set; one molecule per line with a row id. run_records/ - one score file per generator and set (Orthonym engine, openclatura 0.3.2 with self-verification, NISPO 0.1.6), the headline holdout2000 run record with every name and tier, the release-gate verdict, the per-row report on the 1,666-structure reference set, and the single-process timing run on holdout2000. source_data/ - Supplementary Data 1: one CSV per set and generator with every molecule's input SMILES, emitted name, tier, OPSIN read-back, both InChIKeys and the round-trip outcome (18 files, 2,247,728 molecules per generator). scripts/ - the corpus launcher and worker, the runner for the reference generators, the scoring script (OPSIN 2.9.0 with its radical option on, InChIKeys with RDKit 2026.03.5) and the Source Data exporter. README.md and SHA256SUMS - a description of every file and a checksum for every file (verify with sha256sum -c SHA256SUMS). Software: Orthonym engine version 1.0.0 (https://github.com/Steinbeck-Lab/Orthonym, MIT licence); openclatura 0.3.2; NISPO 0.1.6; OPSIN 2.9.0; RDKit 2026.03.5. One scoring script and one scorer setting were used for every generator and every set. Every count row in the paper's tables adds up to the size of its set and can be recomputed from the score files and CSVs here.

提供机构:
Zenodo
创建时间:
2026-09-29
二维码
社区交流群
二维码
科研交流群
商业服务