遇见数据集

Genotype data for a set of 163 worldwide populations

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

Here is a combined dataset of genetic data on 2,643 individuals from 163 worldwide human populations. These genotypes were all generated on Illumina chips (550, 610, 660) for multiple different studies. The two main papers that this dataset was compiled for are: Hellenthal, et al 2014 A Genetic Atlas of Human Admixture History, Science; and Busby, et al 2015 The role of recent admixture in forming the contemporary West Eurasian genomic landscape, Current Biology. The data are in PLINK format and the BusbyWorldwidePopulations.csv file outlines where the different datasets come from. Note that because these two datasets were combined together, not all populations are typed on the same set of SNPs. We have included genotype data on 523,443 SNPs, of which 441,038 are genotyped on at least 97.5% of individuals. Therefore, additional QC steps are required to filter this set down to high quality calls, depending on the subset of samples that are required. Complete information about the populations used is available in the various publications that are outlined in the associated paper. Note that these same populations are available elsewhere and this dataset represents that compiled for the above mentioned papers. UPDATE 11/11/2019 Thanks to some heroic work by Kristján Helgi Swerford Moore at DECODE, I have now updated the population and sample information to more accurately and verbosely label the individuals.

本数据集为整合后的人类遗传数据集,包含来自全球163个人群的2643名个体的基因型数据。所有基因型均通过Illumina芯片(型号涵盖550、610、660)完成分型,数据来源于多项独立研究。本数据集主要为两篇核心论文整合构建:其一为2014年Hellenthal等人发表于《Science》的《人类混血历史遗传图谱》(A Genetic Atlas of Human Admixture History);其二为2015年Busby等人发表于《Current Biology》的《近期混血在当代西欧亚基因组景观形成中的作用》(The role of recent admixture in forming the contemporary West Eurasian genomic landscape)。 数据存储格式为PLINK格式(PLINK format),配套文件BusbyWorldwidePopulations.csv说明了各子数据集的来源。需注意,由于两个原始数据集进行了合并,并非所有人群均使用同一套单核苷酸多态性(Single Nucleotide Polymorphism,简称SNP)位点进行分型。本数据集共包含523443个SNP位点的基因型数据,其中441038个位点可在至少97.5%的个体中成功分型。 因此,根据研究所需的样本子集,需执行额外的质量控制(Quality Control,简称QC)步骤以筛选出高质量的分型结果。关于本数据集所用人群的完整信息,可查阅相关论文中列出的各类已发表文献。 需补充说明的是,上述人群的遗传数据在其他公开渠道亦有提供,本数据集正是为前述两篇论文整合构建的版本。 2019年11月11日更新: 感谢DECODE的Kristján Helgi Swerford Moore所完成的细致工作,现已更新人群与样本的标注信息,以更准确且详尽的方式对个体进行标识。

创建时间:
2021-11-09
二维码
社区交流群
二维码
科研交流群
商业服务