遇见数据集

Metadata for MetaOrion pretraining samples

收藏
Zenodo2026-06-18 更新2026-06-21 收录
官方服务:

资源简介:

To enable large-scale representation learning, we assembled a diverse collection of human metagenome datasets spanning multiple body sites and geographical regions. The final dataset comprises 107,494 samples, with the majority originating from Asia, Europe, and North America, and stool samples accounting for approximately 72%. This dataset provides comprehensive metadata for the cohorts utilized during the MetaOrion pretraining phase, including sample unique identifiers (sample_id), anatomical origins (body_site), precise geographical provenances (country and continent), and original publication identifiers (study_name). To ensure maximal transparency and full community reproducibility, we have curated the public repository footprints—specifically including project accessions (project.accession), raw sequencing run IDs (sequencing.data.accession), and their hosting databases (sequencing.data.db)—alongside essential technical specifications and sequencing attributes, such as total sequencing depth (number_bases), instrument models (sequencing_platform), DNA extraction protocols (DNA_extract_type), and the reference database versions utilized for upstream taxonomic profiling (MetaPhlAn4.database).

提供机构:
Zenodo
创建时间:
2026-06-16
二维码
社区交流群
二维码
科研交流群
商业服务