遇见数据集

SNP genotype matrix for GWAS and Machine Learning analyses

收藏
Zenodo2021-10-25 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>SNP datasets used for GWAS and Machine Learning analyses</strong> All datasets come from the easyGWAS website: https://easygwas.ethz.ch/down/1/ <strong>=== Horton et al. 2012 ===</strong> <strong>1307 Arabidopsis genotypes x 214,057</strong> <strong>SNPs</strong> <strong>1) In the form of a genotype matrix </strong> The file is called Horton2012.raw https://www.nature.com/articles/ng.1042 Preview of the first lines and columns: FID Chr1_657_T Chr1_3102_G Chr1_4648_A Chr1_4880_T Chr1_5975_G Chr1_6063_T Chr1_6449_C<br> 9381 2 2 2 0 0 0 0<br> 9380 0 0 0 0 0 0 2<br> 9378 2 2 2 0 0 0 0<br> 9371 2 2 2 0 0 0 0<br> 9367 0 0 0 2 0 0 0<br> 9363 2 2 2 0 0 0 0<br> 9356 0 2 2 0 0 0 0<br> 9355 2 2 2 0 0 0 0<br> 9354 2 2 2 0 0 0 0 ...etc... PLINK 1.9 was used to convert the .ped and .map file to a .raw format with: <pre><code class="language-bash">plink --file original_data/genotype --recodeA --tab</code></pre> Genotypes are encoded as 0, 1 or 2 with: <pre> SNP SNP_A --- ----- A A -&gt; 0 A C -&gt; 1 C C -&gt; 2 0 0 -&gt; NA </pre> Then only the Family ID was kept (same as individual ID) and other columns (Paternal ID, Maternal ID, Sex, Phenotype) were removed. The corresponding PLINK manual page used is here: https://zzz.bwh.harvard.edu/plink/dataman.shtml#recode <strong>1) In the form of set of files compatible with PLINK out of the box</strong> The archive file is called AtPolyDB_call_method_75_Horton2012.tar.gz and contains three files: genotype.ped: pedigree information from the 1307 ecotypes genotype.map: the SNP positions on the genome phenotypes.pheno: the phenotype value of the 1307 ecotypes

提供机构:
Zenodo
创建时间:
2021-10-25
二维码
社区交流群
二维码
科研交流群
商业服务