遇见数据集

PBsim Taxanomic Classification ID/OOD benchmark dataset

收藏
Zenodo2026-03-17 更新2026-05-26 收录
官方服务:

资源简介:

This dataset was created with the intent to evaluate the performance of fine-tuned genomic language models on both ID and OOD taxanomic classification tasks. The Woltka pipeline was used to compile the full genome dataset, consisting of 4634 genomes, with each genus represented by a single genome. From there, taxa were filtered in to categories with different levels of representation in the training to produce various levels of distribution shift. Reads were then generated using PBsim, to simulate 6kbp PacBio generated reads from the full genomes. There are 5 files available: train.csv - full training dataset of bacterial reads. id_novel_genus.csv - ID test set, if classifying on a family level. Novel genus, but shared family with training ood_novel_family.csv - OOD test set. Novel family, but shared orders ood_nonbacterial.csv - OOD test set. No shared taxanomic lineage with training full_basic_lineage.csv - metadata lineage file

提供机构:
Zenodo
创建时间:
2026-03-17
二维码
社区交流群
二维码
科研交流群
商业服务