Novelty-Tiered Affinity Benchmark (NTAB)
收藏资源简介:
Novelty-Tiered Affinity Benchmark is a protein-ligand binding affinity benchmark derived from ChEMBL 36, designed to minimize data leakage through two complementary strategies: Time split — activities are partitioned by assay publication year (train: before 2022, val: 2022, test: 2023+). Similarity-binned test/val sets — all test and val compounds are labelled by their maximum Morgan Fingerprint (radius 2, 2048-bit) Tanimoto similarity to pre-cutoff compounds, yielding five bins: [0, 0.35), [0.35, 0.5), [0.5, 0.7), [0.7, 1.0), and =1.0. Val and test assays are further filtered to retain only well-characterised (assay, measurement-type) groups: minimum 10 unique compounds, pChEMBL SD ≥ 0.5, equality-relation measurements only, and at most one assay per publication. Files activities.parquet — one row per activity measurement, with split labels, pChEMBL values, compound SMILES, and novelty scores targets.parquet — one row per single-protein target, with UniProt ID, sequence, gene name, and protein classification predictions_*.csv — predictions using the ligand-only baselines as shown in Figure 4 in the preprint Full documentation, column descriptions, inclusion criteria, and code to reproduce the dataset are available at: https://github.com/bamattsson/ntab The preprint is available at: https://www.biorxiv.org/content/10.64898/2026.06.29.735309v1 Attribution This dataset is derived from ChEMBL 36, provided by EMBL-EBI (https://www.ebi.ac.uk/chembl/). ChEMBL is licensed under CC BY-SA 3.0: https://ftp.ebi.ac.uk/pub/databases/chembl/ChEMBLdb/releases/chembl_36/LICENSE Citation If you use this dataset, please cite: https://www.biorxiv.org/content/10.64898/2026.06.29.735309v1 Bibtex: @article{mattsson2026identifying, title = {Identifying and Addressing Systematic Data Leakage in Protein-Ligand Affinity Benchmarks}, author = {Mattsson, Bj{\"o}rn and Walters, W. Patrick}, year = {2026}, journal = {BioRxiv},}



