遇见数据集

SNP datasets and genomes used to benchmark the SNPLift program

收藏
NIAID Data Ecosystem2026-05-01 收录
官方服务:

资源简介:

Motivation: The advent of high-throughput sequencing technologies and availability of reference genomes has provided an unprecedented opportunity to discover and genotype millions of genetic variants in hundreds or even thousands of samples. Variant calling, the identification of genetic variants from raw sequencing data, is a time-consuming and computationally expensive process. Currently, reference genomes are evolving very rapidly and new versions come out more and more frequently. To take advantage of new or improved reference genomes, raw reads alignments, genotype calling, and filtration must typically all be redone. This is a costly and time consuming operation that is not always possible when projects are under time constraints. Results: Here, we present SNPLift, a bioinformatic pipeline that can quickly transfer SNP coordinates from one version of a genome to another, making it possible to rapidly leverage the resources represented by new reference genomes. We tested SNPLift on nine SNP datasets in VCF format from different species (Homo sapiens, Arabidopsis thaliana, Coregonus clupeaformis, Medicato truncatula, Oriza sativa, Salvelinus namaycush, Solanum lycopersicum, Zea mays, and Glycine max). Depending on the species, we accurately lifted between 82.64% and 99.39% of the variants very quickly and efficiently, reducing the required computing power by multiple orders of magnitudes compared to a complete re-analysis using the new genome reference. SNPLift provides an accurate, parallelized, efficient and fast solution to update genome positions, for example for variant calls, based on new reference genomes. Availability and implementation: SNPLift is available at https://github.com/enormandeau/snplift with its documentation and installation procedure. It also contains a script that runs an automated test on a small dataset, composed of 190,443 SNPs in chromosome 1 of Medicago truncatula. SNPLift uses only common tools that are easy to install and works under Linux and MacOS. Methods Nine species are present in the dataset. For each species, two genome versions and one VCF are present. The VCF contains SNPs whose positions refer to the oldest reference genome.

研究背景:高通量测序技术的问世与参考基因组的普及应用,为在数百乃至数千份样本中发掘并完成百万级遗传变异的基因分型提供了前所未有的契机。变异调用(Variant calling)——即从原始测序数据中识别遗传变异的过程——是一项耗时久、计算成本高昂的工作。当前参考基因组的更新迭代极为迅速,新版本的发布频率也日益提升。若要利用新的或优化后的参考基因组,通常需要重新开展原始读段比对、基因型调用与过滤整套流程。此类操作不仅成本高昂且耗时良久,在项目受时间约束的场景下往往难以落地。 研究结果:本文介绍了SNPLift——一款可快速将单核苷酸多态性(Single Nucleotide Polymorphism, SNP)坐标从一个基因组版本转换至另一版本的生物信息学流程,借此可快速依托新参考基因组所承载的研究资源。我们针对来自不同物种的9个VCF格式SNP数据集开展了SNPLift性能测试,涉及物种包括智人(Homo sapiens)、拟南芥(Arabidopsis thaliana)、湖白鲑(Coregonus clupeaformis)、蒺藜苜蓿(Medicago truncatula)、水稻(Oryza sativa)、湖红点鲑(Salvelinus namaycush)、番茄(Solanum lycopersicum)、玉米(Zea mays)以及大豆(Glycine max)。根据物种的差异,该工具可快速高效地完成82.64%至99.39%的变异坐标转换,相较于使用新参考基因组进行完整重分析,所需计算算力降低了多个数量级。SNPLift为基于新参考基因组更新变异位点等基因组位置信息提供了精准、并行化、高效且快捷的解决方案。 可用性与实现方式:SNPLift及其配套文档与安装流程可通过https://github.com/enormandeau/snplift获取。该工具还包含一段可在小型数据集上运行自动化测试的脚本,该数据集包含蒺藜苜蓿1号染色体上的190443个SNP。SNPLift仅依赖易于安装的通用工具,可在Linux与MacOS系统下稳定运行。 研究方法:本数据集涵盖9个物种。每个物种均对应两个基因组版本与一份VCF文件,该VCF文件中的SNP位点坐标均基于较旧的参考基因组。

创建时间:
2023-06-12
二维码
社区交流群
二维码
科研交流群
商业服务