Pre-processed k-mer frequency vectors for Salmonella detection benchmark
收藏资源简介:
All genomic sequences used in this study are publicly available from NCBI GenBank. The dataset comprises 14,156 bacterial genomes: 6,110 Salmonella enterica (serotypes Typhimurium, Enteritidis, Heidelberg, Dublin, Newport, and Infantis) and 8,046 non-Salmonella genomes spanning 12+ bacterial species including Escherichia coli (63 strains), Klebsiella pneumoniae, Shigella flexneri, Pseudomonas aeruginosa, Staphylococcus aureus, Listeria monocytogenes, and others. Pre-processed k-mer frequency vectors, train/validation/test splits (by genome source, 70/15/15%), trained model weights, and complete source code are available at [https://github.com/lalala0721/lnn-salmonella-detection] and [10.5281/zenodo.20597488]. The authors confirm all supporting data, code, and protocols have been provided within the article or through supplementary data files.



