遇见数据集

SNP harmoniser (GRCh37 and GRCh38)

收藏
Zenodo2026-06-24 更新2026-06-28 收录
官方服务:

资源简介:

A comprehensive lookup table for harmonising SNP identifiers across human genome build 37 (GRCh37) and human genome build 38 (GRCh38), including rsID to chr:position mapping Download the lookup table Download the file from this Zenodo archive This file contains all SNPs (MAF >= 0.0001) with a match across human genome build 37 and build 38 (based on NCBI dbSNP 155 patch 13) Matching was done based on matching rsIDs, then using CrossMap This file also contains SNPs in one build but with no match in the other build (relevant columns instead contain "no_match_in_build37" or "no_match_in_build38") How to use the lookup table This file is useful for formatting GWAS summary statistics When matching summary stats to this file it will restrict the sum stats to SNPs only (no indels) that are in builds 37 and 38 of dbSNP 155 patch 13 Matching sumstats to this file allows converting SNP IDs between build 37 and 38 Matching sumstats to this file allows adding rsID if only chromosome and BP were given, and vice versa This file can also be used to create a consistent marker name across multiple sumstats used in a meta-analysis. Thus, all sum stats have the same name for the same SNP allowing correct matching in meta-analysis (regardless of multi-allelic sites) An example script for how to use this lookup table is given in the github repo: 'Example_script_match_GWAS_sumstats_to_dbSNP_ref_file.sh' How the lookup table was created Look at scripts in the github repo for exact details of how this file was made Directory 'dbSNP155_GRCh37.p13' contains all scripts used to download and process the NCBI dbSNP 155 patch 13 file for human genome build 37 Directory 'dbSNP155_GRCh38.p13' contains all scripts used to download and process the NCBI dbSNP 155 patch 13 file for human genome build 38 Directory 'dbSNP155_GRCh37.p13_vs_dbSNP155_GRCh38.p13' contains all scripts used to match SNPs across build 37 and build 38 Important information to note SNPs were matched across human genome builds using rsIDs then CrossMap This was done, rather than just using CrossMap, because it is the reccommended way of converting SNPs between builds See FAQ 'Converting SNPs between assembly versions' at https://genome.ucsc.edu/FAQ/FAQreleases.html for more info: 'While the LiftOver tool exists to convert coordinates between assemblies, it is NOT recommended to use LiftOver to convert SNPs between two assembly versions. Using LiftOver to convert a very small region, especially a single base SNP, is not always a trivial task as the alignment may not get a high enough score to be considered a successful conversion.' This file only contains SNPs with MAF >= 0.0001 If all SNPs from dbSNP were included this would be > 1 billion making the memory and time to do this processing not possible So I filtered to only use MAF > 0.0001 for convenience However, most GWAS are QCed to MAF > 0.01. Thus, this dbSNP refrence file should have all of the SNPs in your GWAS cohort This file only contains SNPs Indels were removed

提供机构:
Zenodo
创建时间:
2026-06-24
二维码
社区交流群
二维码
科研交流群
商业服务