遇见数据集

1000 Genomes Project hg38, phased, biallelic, normalized

收藏
Zenodo2025-08-22 更新2026-05-26 收录
官方服务:

资源简介:

Pre-processed genotypes from the 1000 Genomes Project. Starting from the per-chromosome phased VCFs found here: left-align & normalize against hg38, atomize, and split multiallelics using bcftools concatenate all chromosomes together convert to PLINK 2 format, compressing the PVAR file with zstd to be compatible with PLINK 2 compress the PGEN file with xz PLINK -> genoray SVAR -> tar archive -> xz compress Note that the SVAR file is ~90 GB decompressed and must be decompressed for use. 32 GB of RAM is recommended for decompression to avoid an out-of-memory error.

本数据集为千人基因组计划(1000 Genomes Project)的预处理基因型数据,原始数据来源于下述链接中提供的每条染色体相位化变异识别格式(VCF)文件: 使用bcftools对原始数据执行左对齐、以hg38参考基因组为基准完成标准化、原子化处理并拆分多等位基因位点;随后将所有染色体的数据集拼接合并。 将数据转换为PLINK 2格式,同时使用zstd压缩PVAR文件以适配PLINK 2工具,再使用xz压缩PGEN文件。 通过PLINK工具生成genoray格式的SVAR文件,将其打包为tar归档文件后,使用xz完成最终压缩。 需注意:该SVAR文件解压缩后体积约为90 GB,使用前必须完成解压缩;为避免内存溢出错误,解压缩时建议配置32 GB及以上内存。

提供机构:
Zenodo
创建时间:
2025-08-22
二维码
社区交流群
二维码
科研交流群
商业服务