遇见数据集

Ames mutagenicity curated dataset - TDC (ChemPharos datasets)

收藏
Zenodo2026-03-10 更新2026-05-26 收录
官方服务:

资源简介:

A curated and enriched Ames Mutagenicity dataset of chemical compounds. The data were initially collected by Xu et al. (https://doi.org/10.1021/ci300400a) and retrieved from the Therapeutics Data Commons (TDC) platform (https://tdcommons.ai/single_pred_tasks/tox/#ames-mutagenicity). TDC provides Python functions for data splitting to facilitate the training of machine learning models. In this case, the TDC protocol is employed for random splitting of the data into training (70%), validation (10%), and test (20%) subsets, as indicated in the “Subset” column (column ALY). To ensure that the mutagenicity endpoint corresponds precisely to a single structure rather than an ambiguous spatial arrangement or a racemate, compounds with ambiguous stereochemistry are filtered out. Duplicate entries, both within and across subsets, are also removed. This results in a final dataset comprising 3724 training compounds, 549 validation compounds, and 1089 test compounds. Each compound is labelled (“Y”-column C) as either mutagenic/Ames positive (class 1) or non-mutagenic/Ames negative (class 0). Compounds in all subsets are sanitized, standardized, and transformed into canonical SMILES using RDKit to ensure consistency and eliminate representational redundancies (column B). The dataset is enriched with 777 Mold2 (columns D -ACZ), 210 2D RDKit (columns ADA-ALB), and 11 3D RDKit descriptors (columns ALC-ALM). For the compounds for which no conformers are found, 3D descriptors are not computed. Whenever available, additional identifiers such as PubChem CIDs, and InChIKeys are included, retrieved from the PubChem database using the compounds’ Canonical SMILES, by employing the Enalos+ KNIME nodes (columns ALN-ALX). More curated datasets are available via chemPharos: https://db.chempharos.eu/datasets/Datasets.zul

提供机构:
Zenodo
创建时间:
2026-03-10
二维码
社区交流群
二维码
科研交流群
商业服务