遇见数据集

GenPept-Curated-2025: An annotation-derived, high-identity-controlled benchmark for antimicrobial peptide annotation prediction

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

Overview GenPept-Curated-2025 version 1.1 is an annotation-derived, sequence-level benchmark of 11,000 unique GenPept/NCBI Protein sequences for binary antimicrobial-peptide (AMP) annotation prediction. Scientific purpose The dataset supports reproducible evaluation of machine-learning methods that predict AMP annotation from peptide or short protein sequences. Its balanced benchmark design and fixed partitions provide a common basis for method comparison. Data source Sequences were retrieved from GenPept/NCBI Protein molecular sequence resources, covering Bacteria, Archaea and Fungi and sequence lengths of 10–200 amino-acid residues. Preserved operational NCBI query specifications for the main and precursor-related source queries are included in the package. Dataset construction The documented construction workflow uses IPG-based deduplication, annotation-based labeling, and CD-HIT clustering at 90% pairwise sequence identity, followed by cluster-intact partitioning and post-split MMseqs2 verification. The released canonical CSV contains 11,000 distinct sequences. In this release, sequence_clean equals sequence and len_clean equals length for every record. The public release preserves the final canonical dataset and direct split subsets; it does not contain the full analysis-code repository. Class definition AMP identifies records with qualifying AMP annotations under the documented operational labeling rules. The non-AMP label identifies annotation-negative/unlabeled comparison records: absence of evidence is not evidence of absence. These comparison records are not experimentally confirmed inactive peptides. Labels are annotation-derived rather than uniform assay-defined outcomes. Dataset composition There are 11,000 records: 5,500 AMP annotation-positive records and 5,500 annotation-negative/unlabeled comparison records. There are 11,000 unique sequences and 11,000 unique sample IDs. The 50:50 composition is a controlled benchmark design and is not an estimate of natural AMP prevalence. Sequence lengths range from 10 to 200 amino-acid residues. Train/validation/test partition The authoritative split field assigns 7,700 records to train (3,850 per class), 990 to validation, encoded as val (495 per class), and 2,310 to test (1,155 per class). The convenience files data/splits/train.csv, data/splits/validation.csv and data/splits/test.csv are exact direct subsets of the canonical split field and do not define alternative partitions. High-identity and component leakage control Partition control is component-aware: the 10,666 released operational connected components, identified by comp_id, do not cross the canonical partitions. The preserved MMseqs2 audit summary reports zero cross-partition matches under its reported settings of at least 90% identity and at least 0.8 coverage of the shorter sequence. This is a report of the preserved audit, not a newly executed sequence-similarity search. CD-HIT and MMseqs2 operational screens have different implementations and their definitions are not interchangeable. High-identity control does not establish universal remote-homology independence, biological-family independence, or absence of every possible source of information leakage. comp_id is an operational partition-control identifier, not a biological-family annotation. Canonical dataset The authoritative file is data/GenPept_Curated_2025_primary_split_v1.1.csv. Its SHA-256 is 421e55265e1462052f633c961f8d8cc20ce5e510284ffc36afc6e29a40b79b6c. The release version is 1.1. Older dataset.csv, balanced_11000.csv or legacy split artifacts must not be substituted for this canonical public dataset. Package contents The ZIP contains the canonical CSV; the three direct split CSVs; README.md; DATA_DICTIONARY.csv; SOURCE_AND_RIGHTS.md; VERSION.txt; CHANGELOG.md; ZENODO_METADATA.md; provenance/query_specification_from_preserved_rules.csv; validation/PUBLIC_RELEASE_DATASET_QA.json; validation/canonical_integrity.json; validation/ZENODO_DATASET_VALIDATION.json; validate_dataset.py; MANIFEST.csv; and SHA256SUMS.txt. Intended use Use train for model fitting, val for model selection and validation, and the fixed test partition for final evaluation. Report preprocessing, sequence representation, model hyperparameters, random seeds, decision-threshold selection and evaluation metrics. Retain the released split for direct benchmark comparisons. The dataset supports research on sequence-based annotation prediction; it does not establish experimental antimicrobial activity for individual sequences. Limitations Annotation-derived labels may reflect the source annotation process. The comparison class is not an experimentally validated inactive class. Sequence length, source annotation and taxonomic structure can carry predictive information. The canonical test set is an in-benchmark holdout with operational high-identity control; it is not a prospective, temporal, taxonomic, biological-family-held-out or clinical cohort. The canonical public CSV does not retain accession, taxonomy, original annotation text, IPG, precursor or other source metadata absent from its documented columns. Row-level source provenance cannot be reconstructed from fields that are not present. Reproducibility information The package includes the preserved executable NCBI query specifications, a column dictionary, QA and integrity summaries, a manifest with file sizes and SHA-256 hashes, package checksums, and a portable Python standard-library validator. Run python validate_dataset.py from the extracted release directory to check canonical identity, record and class counts, sequence properties, component separation and exact split subsets. The provided validator does not rerun source retrieval, CD-HIT or MMseqs2. Distinguish reproducibility of the documented retrieval and curation rules from row-level provenance available in the public CSV. Cite the Zenodo DOI of the exact dataset version used; any associated article DOI is a separate related identifier. License and source rights Creative Commons Attribution 4.0 International (CC BY 4.0; SPDX: CC-BY-4.0) applies to creator-owned curation, operational labels and derived annotations, split assignments, documentation, validation metadata and compilation, only to the extent the creators hold the relevant copyright or database rights. This statement does not claim ownership of, or purport to relicense, third-party rights in underlying NCBI/GenBank sequence records. As explained in SOURCE_AND_RIGHTS.md, NCBI states that it places no restrictions on use or distribution of molecular data in its databases, while original submitters or jurisdictions may claim patent, copyright or other rights that NCBI cannot grant. Users remain responsible for respecting applicable source-record rights. See https://www.ncbi.nlm.nih.gov/home/about/policies/, https://www.ncbi.nlm.nih.gov/genbank/about/ and https://creativecommons.org/licenses/by/4.0/.

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务