MAOMAO: An Ontology-Guided FAIR Resource for Harmonized Peptide Toxicity Data
收藏资源简介:
MAOMAO (*Metadata-Aware Ontology for Multi-source Annotation Organization*) is an ontology-guided, provenance-aware, and uncertainty-aware resource designed to support the harmonization, interpretation, and reuse of peptide toxicity data. It addresses the fragmentation of existing peptide datasets, which frequently differ in terminology, annotation criteria, evidence quality, class definitions, and reporting practices. By integrating heterogeneous sources under a controlled toxicity vocabulary and an explicit evidence model, MAOMAO provides a common framework for representing peptide toxicity annotations without treating missing, negative, unlabeled, and conflicting information as equivalent. The resource integrates data from 54 databases, published datasets, repositories, supplementary materials, and prediction resources. After sequence validation, harmonization, and endpoint-level evidence integration, the final resource contains 71,701 unique peptide sequences annotated across seven controlled toxicity endpoints: toxic, cytotoxic, hemolytic, cytolysis, neurotoxic, embryotoxic, and ichthyotoxic. Each sequence–endpoint combination is represented using one of five mutually exclusive evidence states: positive, negative, ambiguous, unlabeled, or no information. Positive evidence is propagated through an explicit toxicity hierarchy, while endpoint-specific ambiguity is preserved and negative evidence remains endpoint-specific. The Zenodo deposit is organized into five complementary layers. The Core Layer contains the harmonized sequence-level dataset, source- and endpoint-level processed data, ambiguity-support records, hierarchy audits, and provenance metadata. The Descriptor Layer provides 41 physicochemical and sequence-derived descriptors for the complete peptide collection. The Embedding Layer contains ten precomputed protein language model representations and one one-hot baseline. The Benchmark Layer provides reproducible train–validation–test partitions for four binary toxicity endpoints using two partitioning strategies, 30 random seeds, and five folds. These partitions can be combined with any of the 11 numerical representations to reconstruct 2,640 benchmark configurations. The Documentation Layer provides the controlled vocabulary, metadata schema, licensing guidance, version records, release inventories, and checksums. MAOMAO is intended to support transparent toxic-peptide data curation, exploratory analysis, representation comparison, reproducible machine-learning evaluation, and the development of toxicity-aware peptide prediction and design workflows. The deposited files preserve the correspondence among sequences, endpoint states, numerical representations, partitions, processing parameters, and provenance records, enabling users to inspect and reuse the resource at different levels of detail. For detailed information about file organization, evidence-state encoding, hierarchy rules, metadata fields, licensing, and recommended usage, see the included README.md and the documentation_layer directory.



