THEOBROMA v1.35: an open multi-kingdom natural-products database
收藏资源简介:
THEOBROMA is an aggregated open natural-products database containing 1,132,805 compounds from 29 source databases across four biological kingdoms plus an unresolved category. The corpus preserves full 27-character InChIKeys and exposes the resulting 486,032 connectivity families, which group stereoisomers togetherwith protonation and isotopic variants, as a directly queryable resource. Provenance is tracked at compound granularity for both classification (source- inherited, NPClassifier tool, or model-inferred) and licensing (per-compound resolution across all attesting sources under a most-restrictive-wins rule, with a least-restrictive bound reported alongside it). Open-licensed compound coverage is 79.5% of the corpus under the most restrictive resolution and 91.2% under the least restrictive. License tiers are recorded per compound in the license_tier and tier_rank_min columns for downstream filtering: 891,860 compounds resolve to CC BY 4.0, 8,243 to CC0, 225,536 to CC BY-NC 4.0, and 7,166 to sources with unspecified terms. The deposit comprises six files: a PostgreSQL custom-format archive of 18 relational tables (compounds, compound_taxonomy, resolved_taxonomy, per_source_license_attestation, source_license_ref, compound_synonyms, compound_region_map, admet, scaffolds, the NPClassifier ontology tables, and the taxonomy reference tables) with its checksum, the classifier reproducibility bundle (XGBoost model, PCA basis, Optuna study, per-class evaluation, threshold calibration), the per-source reproducibility manifest sources.yaml, a README giving restore and verification commands, and SHA256SUMS covering all files. Restore requires the pg_trgm extension for the trigram name index. This record declares three licenses because the corpus is not uniformly licensed. CC BY 4.0 covers the contributed layer: the license audit and resolved tiers, the classification, the taxonomy resolution, and the schema. CC BY-NC 4.0 and CC0 1.0 are declared because upstream source terms flow through unchanged at compound granularity. Downstream use should be governed by the license_tier column rather than by any single record-level designation. The accompanying live database is at https://theobroma.l3s.uni-hannover.de and source code at https://github.com/ThorKlm/theobroma. This deposit accompanies the preprint at https://doi.org/10.64898/2026.06.12.731585. A Parquet mirror of the principal tables, with a dataset viewer and load_dataset() support, is available at https://huggingface.co/datasets/ThorKl/theobroma. Supersedes the 12 August 2026 deposit of v1.35. Corrections to the corpus: kingdom assignment for 6,115 compounds where the majority vote contradicted the resolved lineage; separator normalisation in effective_pathway and effective_superclass; descriptor computation for 68,395 compounds whose source records had been skipped. Corrections to sources.yaml: citation DOIs and source URLs updated, a YAML parsing problem resolved, and the enrichment and pipeline blocks, which had retained v1.34 classifier values and stale software versions. Classifier values and the Optuna version are now taken from the artefacts in classifier_reproducibility.zip.



