遇见数据集

K-FluDB

收藏
Zenodo2025-07-15 更新2026-05-26 收录
官方服务:

资源简介:

K-FluDB: A Novel K-Mer Based Database for Enhanced Genomic Surveillance of Influenza A Viruses K-FluDB is a compressed database composed of distinct sub-sequences specific to 50 influenza A subtypes. It includes unique sequences for all 18 hemagglutinin (HA) and 11 neuraminidase (NA) subtypes. The original influenza sequences were obtained from the NCBI database on May 8, 2022, comprising a total of 895,900 Influenza A sequences. To generate this database, sequences were first subsampled based on genomic segment and variant group, resulting in 81,262 sequences. These sequences served as input for the PanGen-InfluenzaA tool (GitHub), which constructs pangenomes by identifying both unique subtype-specific sequences and sequences shared across multiple subtypes. Repository Contents This repository contains a ZIP archive with three folders, each corresponding to pangenome datasets designed for reads of 74, 150, and 300 nucleotides in length. Each folder includes the following files: Dispensable data files (1_dispensable.fasta to 8_dispensable.fasta):These files contain the dispensable genomic fragments for each of the eight segments of the Influenza A virus. Unique data files (1_unique.fasta to 8_unique.fasta):These files contain the unique genomic fragments for each segment. The recommended files for mapping against genomic reads are 4_unique.fasta and 6_unique.fasta, corresponding to segments 4 and 6, which are the targets commonly used for Influenza A subtyping. pangenome.fasta:This file contains the combined set of unique and dispensable genomic fragments for all eight segments. unique.fasta:This file includes only the unique genomic fragments for all eight segments. Compression Efficiency and Classification Accuracy K-FluDB achieves a relative compression index of 96.54% when using the complete pangenome and 99.64% when considering only subtype-specific sequences. The average precision for correctly classifying Hx and Nx subtypes using the unique sequences is 99.2% and 99.71%, respectively. This database provides a highly efficient and accurate resource for influenza A subtype classification while significantly reducing the storage and computational requirements associated with full-genome analyses. Acknowledgements This research was partially supported by grants by PAPIIT-DGAPA-IN230523 awarded to BT. The first author gratefully acknowledges the scholarship provided by CONAHCYT. We also extend our sincere appreciation to the National Autonomous University of Mexico (UNAM) for granting access to the MIZTLI supercomputer, supported by the General Directorate of Computing and Information and Communication Technologies (DGTIC) through project LANCAD-UNAM-DGTIC-350.

提供机构:
Zenodo
创建时间:
2025-07-15
二维码
社区交流群
二维码
科研交流群
商业服务