遇见数据集

Relate Book Chapter example files

收藏
Zenodo2025-08-15 更新2026-05-26 收录
官方服务:

资源简介:

Overview This repository provides the data to run the practical examples displayed in Population genomics by reconstructing whole genome genealogies using Relate. This chapter serves as a practical user guide for Relate, a method that enables large-scale whole-genome genealogy inference. Relate reconstructs coalescent trees across the genome and provides downstream tools for a broad range of population genetic analyses. Here, we provide an example dataset including both modern and ancient human genomes to demonstrate key applications. Note: This dataset contains the same files as in the following three Zenodo records: Twigstats scripts and example dataset Relate input files Relate example for quantifying positive selection The only difference is that certain populations have been merged into one single population in the poplabel file to simplify the interpretation of results. Otherwise, the content is unchanged—only reorganized here into a single, structured record for clarity. Contents 1. Example Data for Relate & Twigstats on Imputed Ancient Genomes Self-contained dataset for testing Relate and Twigstats pipelines on ancient genomes imputed with GLIMPSE. Includes GLIMPSE-imputed ancient VCFs, modern reference VCFs (e.g., 1000 Genomes Project subset), population label files, sample ages, and required Relate input files. Provided for both: Chromosome 1 (example_chr1.tgz) Whole genome (example_wg.tgz) For the whole genome example, Relate_input_files.tgz is also required as it contains the Relate input resources for hg37 and hg38, including ancestral genome sequences, genome accessibility masks, Recombination maps, and the human-wide coalescence rates estimated from 1000 Genomes and SGDP datasets, usable as priors in Relate. If you choose the whole-genome option, please make sure to move Relate_input_files inside of example_wg. In this example, we are using data from the 1000 Genomes Project dataset (Nature 2015). We additionally use low coverage shotgun genomes from Anglo-Saxon contexts, British Iron/Roman Age, Irish Bronze Age, and the Scandinavian Early Iron Age (Cassidy et al, PNAS 2016; Martiniano et al, Nature Communications 2016; Anastasiadou et al, Communications Biology 2023; Schiffels et al Nature Communications 2016; Gretzinger et al Nature 2022; Rodriguez-Varela et al Cell 2023). All ancient genomes were imputed using GLIMPSE. 2. Example Data for Quantifying Positive Selection Using Relate A fully contained example of running Relate on CEU samples of the 1000 Genomes Project dataset (Nature 2015). We include data from a few Mb around the LCT gene and construct Relate trees, and subsequently compute selection p-values using the module RelateSelection. Usage Notes For the first part of the chapter, download either: example_chr1.tgz (chromosome 1 only), or example_wg.tgz plus Relate_input_files.tgz (whole genome + ancestral genome, recombination map, and mask files for GRCh37 and GRCh38 in Relate-compatible formats) For training purposes, we recommend using the chromosome 1 example—it is smaller and runs faster. For the selection example, download the corresponding archive (relate_1000G_LCT.tgz). This file is self-contained. Both the first and second parts of the chapter can be run independently. In any case, please extract tar balls using tar -xzvf filename.tgz. Scripts The scripts required to run the analyses are available on the associated GitHub repository https://github.com/ainacolovila/BookChapter_Relate. The GitHub repository provides step-by-step run.sh pipelines for: Preparing VCFs Creating Relate input files Running Relate genealogies Performing admixture analysis (Twigstats) Estimating population sizes Conducting MDS Computing selection p-values Dependencies Please make sure to have the following softwares and packages intalled: bcftools from https://samtools.github.io/bcftools/howtos/install.html (make sure that BCFTOOLS_PLUGINS is set to the correct plugin path). Relate Twigstats Recommended R packages: relater, ggplot2, dplyr, tidyr, plyr, umap

提供机构:
Zenodo
创建时间:
2025-08-15
二维码
社区交流群
二维码
科研交流群
商业服务