Curated carbapenemase protein sequence dataset and reproducibility files
收藏资源简介:
This dataset contains curated protein-sequence data and reproducibility materials supporting an integrated analysis of carbapenemase annotation confidence, exact-sequence redundancy, phylogenetic family/lineage resolution, taxonomic representation, and temporal database representation. The source collection comprised 4,178 protein accession records, dereplicated into 3,349 exact unique amino-acid sequences. Integrated validation using CARD RGI, NCBI AMRFinderPlus, Pfam-domain evidence, class-specific catalytic motifs, and phylogenetic analyses retained 3,246 unique proteins representing 4,069 accession records. The primary analytical cohort comprised 2,548 full-length QC-passing unique proteins representing 3,294 accession records, while the supported secondary cohort comprised 698 unique proteins representing 775 accession records. The deposited archive includes accession-to-sequence mappings, exact-sequence clusters, cohort assignments, integrated annotation and validation outputs, final family/lineage assignments, phylogenetic alignments and trees, taxonomic and temporal analysis tables, Supplementary Tables S1–S20, analysis scripts, and reproducibility records. Source protein records remain traceable to their public UniProtKB and UniParc accessions. Taxonomic counts describe representation within the curated dataset and should not be interpreted as prevalence, transmission frequency, or host specificity. Temporal variables correspond to database-record creation/deposition dates and should not be interpreted as isolate-collection dates, biological emergence, or discovery dates.



