遇见数据集

A Six-Class Semi-Supervised Dataset for Drug Resistance Classification from Antimicrobial DNA Sequences

收藏
Zenodo2026-07-30 更新2026-08-01 收录
官方服务:

资源简介:

This dataset is a six-class, semi-supervised derivative of the dataset “A Dataset for Drug Resistance Classification from Antimicrobial DNA Sequences.” It provides predefined labeled, unlabeled, validation, and test splits for developing and evaluating semi-supervised sequence classification models for antimicrobial resistance (AMR). The dataset contains curated antimicrobial resistance gene sequences originating from the Comprehensive Antibiotic Resistance Database (CARD) and MEGARes v3.0. Resistance annotations were standardized using the Antibiotic Resistance Ontology (ARO), including information on Drug Class, Resistance Mechanism, and Gene Family where available. The derivative dataset focuses on six major antimicrobial drug classes: Beta-lactams Aminoglycosides Fluoroquinolones Tetracyclines Glycopeptides MLS (Macrolide-Lincosamide-Streptogramin) The dataset contains a total of 3,730 nucleotide sequences and is divided into four predefined subsets: Labeled training set: 300 sequences Unlabeled training set: 2,481 sequences Validation set: 192 sequences Test set: 757 sequences The labeled training subset is class-balanced and contains 50 sequences for each of the six drug classes. The unlabeled training subset contains nucleotide sequences without class or annotation columns and is intended for unsupervised or semi-supervised representation learning. The validation and test sets retain their Drug Class labels and associated AMR annotations for model selection and final evaluation. For labeled samples, the dataset includes the following fields where available: ARO identifier Full-length nucleotide sequence Gene Family Resistance Mechanism Drug Class MEGARes identifier MEGARes header Original source header Numerical class label The represented resistance mechanisms include antibiotic inactivation, antibiotic target alteration, antibiotic efflux, antibiotic target protection, and reduced permeability to antibiotics. Gene family annotations include beta-lactamases, aminoglycoside-modifying enzymes, efflux-associated proteins, ribosomal protection proteins, rRNA methyltransferases, and other AMR-related families. Full-length nucleotide sequences are provided. Researchers may truncate or tokenize the sequences according to the requirements of their selected model. The dataset is suitable for supervised and semi-supervised AMR classification, sequence representation learning, genomic foundation-model evaluation, and benchmarking under limited-label conditions. This dataset does not represent a replacement for the original nine-class dataset. It is a derived version with a reduced six-class label space and predefined labeled and unlabeled subsets specifically designed for semi-supervised learning experiments. Users should cite both this derivative dataset and the original dataset from which it was constructed.

提供机构:
Zenodo
创建时间:
2026-07-30
二维码
社区交流群
二维码
科研交流群
商业服务