遇见数据集

Data for "A Hybrid Pipeline for Feature Reduction, and Ordinal Classification to Predict Antimicrobial Resistance from Genetic Profiles"

收藏
Zenodo2026-04-27 更新2026-05-26 收录
官方服务:

资源简介:

Notation Rules for Variable Interpretation We provide below the notation system used to encode gene-level features in the dataset: {P,N}{A,P}#{R,P}#_{A,P}X#Y The notation is divided by underscores (_). The symbol # represents a number composed of at least one digit. Curly braces {} indicate allowed options. For example, {P,N} means either P or N must appear in that position. Each entry consists of three or four components. 1. Genomic Location Indicates whether the gene is located on a plasmid. P – Located on a plasmid N – Not located on a plasmid 2. Gene Function Indicates the functional annotation source. A – Function annotated in CARD P – Function identified by Prokka # – Function identifier When the function type is A, # corresponds to a CARD ARO identifier. When the function type is P, # corresponds to an autogenerated identifier that maps to a standardized functional annotation index provided in a companion file. For example: P0 → hypothetical protein P1 → tRNA U34 carboxymethyltransferase P2 → Cytosol non-specific dipeptidase P3 → Proline/betaine transporter P4 → ABC transporter glutamine-binding protein GlnH P5 → Hippurate hydrolase ... The complete mapping between Prokka-derived identifiers (P#) and their functional descriptions is provided in the accompanying file: Prokka_Function_Reference_Table.txt 3. Molecule Type and Family Indicates the type of gene product. R – The gene produces an RNA P – The gene produces a protein # – Autogenerated identifier, unique within each function category 4. Optional Mutation Annotation Indicates the presence of a mutation. A – Mutation annotated in CARD P – Mutation detected through gene family alignment X – Reference nucleotide or amino acid For P-type mutations, the reference corresponds to the most common nucleotide or amino acid within the alignment (or the first in alphabetical order in case of ties). For A-type mutations, the reference corresponds to a CARD gene from a susceptible strain. If X = "-", this indicates an insertion relative to the reference. # – Position of the mutation in the RNA or protein sequence Y – New nucleotide or amino acid If Y = "-", this indicates a deletion relative to the reference. Annotation Priority Rules CARD annotations (functions and mutations) take precedence over Prokka and alignment-derived annotations, respectively. CARD annotations replace Prokka or alignment annotations when all of the following conditions are met: Same contig Same strand Same gene type (RNA or protein) Overlap greater than 95% Counting Rules Gene counts and mutation counts are reported in separate columns. Each column represents counts from a single gene family. A gene family may appear across multiple columns. Genes annotated as having unknown function are not guaranteed to be grouped consistently. Example P_A3008823_P0 A gene located on a plasmid that produces a protein with CARD ARO identifier 3008823 (PC1 protein). Software and Annotation Pipeline All genomic features and annotations were generated using standardized bioinformatic tools as follows: Prokka was used for gene prediction and primary genome annotation (CDS and functional annotation). RGI (Resistance Gene Identifier) was used to identify antimicrobial resistance genes and resistance-associated SNPs based on CARD. Barrnap was used to detect ribosomal RNA (rRNA) genes. Infernal was used to identify non-coding RNAs (ncRNAs). Platon was used to determine plasmid origin (plasmid vs chromosomal classification). Sourmash was used for genome-level taxonomic classification. All tools were executed using consistent parameter settings across genomes.

提供机构:
Zenodo
创建时间:
2025-06-09
二维码
社区交流群
二维码
科研交流群
商业服务