遇见数据集

oriGen 1,427 WGS - PCA and ADMIXTURE Reference Datasets for Projection

收藏
Zenodo2026-07-17 更新2026-08-02 收录
官方服务:

资源简介:

Projection Reference Datasets This directory contains the necessary files to project external genomic data into the ancestry and PCA space defined in this study. These files allow researchers to utilize our specific population genomic context without requiring access to the raw genotype data. Files Included PCA_weight.tsv.gz SNP weights (loadings) for the first 10 Principal Components. references_for_admixture.5.P.gz Ancestral allele frequencies for ADMIXTURE (K=5). variant_list_for_admx_and_PCA.tsv.gz The reference variant map used for both analyses. ##### The PCA_weight.tsv.gz file contains the loadings generated by smartpca (EIGENSOFT). Columns: chromosome, and physical position (bp). The subsequent 10 columns correspond to the weights for PC1 through PC1 Usage: These weights can be used with the projectpca tool from the EIGENSOFT suite or via manual matrix multiplication to place new samples into our PCA space. The references_for_admixture.5.P.gz file contains the P matrix (ancestral allele frequencies) for K=5. Format: 1,657,644 rows (one per SNP) and 5 columns (one per ancestral component). Note: Ensure your input file has been filtered to match the variants in our variant_list_for_admx_and_PCA.tsv.gz. The variant_list_for_admx_and_PCA.tsv.gz is the "key" for both the PCA weights and ADMIXTURE frequencies. It contains 1,657,644 lines in the exact order they appear in the analysis files. Important Note on Strand Alignment: Before projecting data, please ensure that your alleles are aligned with the ALT and REF columns provided here. If your data uses the opposite strand for a specific SNP, you must flip the genotype (0 becoming 2, 2 becoming 0) or the sign of the PCA weight to ensure accurate projection.

提供机构:
Zenodo
创建时间:
2026-07-17
二维码
社区交流群
二维码
科研交流群
商业服务