Predicted genome-wide chromatin contact differences among 71 bonobos and chimpanzees
收藏资源简介:
This file contains predicted chromatin contact differences in HFF cells using Akita among pairs of 71 bonobos and chimpanzees at 4,420 ~ 1 Mb genomic windows in the panTro6 genome. Each entry corresponds to a pairwise comparison at a given window. Data per comparison includes the individual IDs in the pairwise comparison, lineages represented, chromosome, position, window ID, mean squared error, Spearman correlation, divergence (1 - Spearman correlation), and the number of nucleotide differences for the pair at the given window., All data in this file were generated using publicly available datasets (see below). Variants were filtered using bcftools to retain high-quality sites and high-quality genotypes. The dataframe was created using Pandas in a Python Jupyter notebook. Data used to generate these predictions: chimpanzee reference sequence (https://hgdownload.soe.ucsc.edu/goldenPath/panTro6/bigZips/) Illumina short reads for 71 bonobos and chimpanzees (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA189439, https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB15086) , , This file contains descriptions of the column headers for \"HFF\_comparisons.txt\". * ind1 = individual 1 in comparison * ind2 = individual 2 in comparison * lineages = lineages represented in comparison (ppn = bonobo, pte = Nigeria-Cameroon chimpanzee, pts = eastern chimpanzee, ptt = central chimpanzee, ptv = western chimpanzee) * chr = chromosome * window\_start = 0-based position in panTro6 coordinates * window = window ID * mse = mean squared error between contact maps for individual 1 and individual 2 * spearman = Spearman correlation between contact maps for individual 1 and individual 2 * divergence = 1 - Spearman correlation between contact maps for individual 1 and individual 2 * seq\_diff = the number of nucleotide differences between individual 1 and individual 2 for the window
本文件包含基于Akita模型,在panTro6参考基因组的4420个约1Mb的基因组窗口中,针对71只倭黑猩猩与黑猩猩的配对样本,预测得到的HFF细胞(人包皮成纤维细胞)内染色质接触差异数据。每条记录对应单个指定窗口下的配对比较结果。单次比较的数据集包含:配对比较中的个体ID、所代表的演化支、染色体、位点、窗口ID、均方误差(mean squared error)、斯皮尔曼相关性(Spearman correlation)、分歧度(1 - 斯皮尔曼相关性),以及该配对在指定窗口内的核苷酸差异数量。 本文件内所有数据均基于公开数据集生成(详见下文)。变异位点通过bcftools进行过滤,以保留高质量位点与高质量基因型。该数据框通过Python环境下的Jupyter Notebook(Jupyter笔记本)与Pandas库构建完成。 用于生成上述预测的数据集如下: 黑猩猩参考序列(https://hgdownload.soe.ucsc.edu/goldenPath/panTro6/bigZips/) 71只倭黑猩猩与黑猩猩的Illumina短读长测序数据(https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJNA189439, https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB15086) 本文件包含对"HFF_comparisons.txt"列标题的说明。 * ind1:比较中的个体1 * ind2:比较中的个体2 * lineages:比较中涉及的演化支(ppn=倭黑猩猩(bonobo),pte=尼日利亚-喀麦隆黑猩猩,pts=东部黑猩猩,ptt=中部黑猩猩,ptv=西部黑猩猩) * chr:染色体 * window_start:panTro6坐标系下的0起始位点 * window:窗口ID * mse:个体1与个体2的染色质接触图谱间的均方误差 * spearman:个体1与个体2的染色质接触图谱间的斯皮尔曼相关性 * divergence:个体1与个体2的染色质接触图谱间的分歧度(即1 - 斯皮尔曼相关性) * seq_diff:该窗口内个体1与个体2间的核苷酸差异数量



