遇见数据集

4C-ker: A method to reproducibly identify genome-wide interactions captured by 4C-Seq experiments

收藏
官方服务:

资源简介:

4C-Seq has proven to be a powerful technique to identify genome-wide interactions with a single locus of interest (or "bait") that can be important for gene regulation. However, analysis of 4C-Seq data is complicated by the many biases inherent to the technique. An important consideration when dealing with 4C-Seq data is the differences in resolution of signal across the genome that result from differences in 3D distance separation from the bait. This leads to the highest signal in the region immediately surrounding the bait and increasingly lower signals in far-cis and trans. Another important aspect of 4C-Seq experiments is the resolution, which is greatly influenced by the choice of restriction enzyme and the frequency at which it can cut the genome. Thus, it is important that a 4C-Seq analysis method is flexible enough to analyze data generated using different enzymes and to identify interactions across the entire genome. Current methods for 4C-Seq analysis only identify interactions in regions near the bait or in regions located in far-cis and trans, but no method comprehensively analyzes 4C signals of different length scales. In addition, some methods also fail in experiments where chromatin fragments are generated using frequent cutter restriction enzymes. Here, we describe 4C-ker, a Hidden-Markov Model based pipeline that identifies regions throughout the genome that interact with the 4C bait locus. In addition, we incorporate methods for the identification of differential interactions in multiple 4C-seq datasets collected from different genotypes or experimental conditions. Adaptive window sizes are used to correct for differences in signal coverage in near-bait regions, far-cis and trans chromosomes. Using several datasets, we demonstrate that 4C-ker outperforms all existing 4C-Seq pipelines in its ability to reproducibly identify interaction domains at all genomic ranges with different resolution enzymes. 4C-Seq experiments from Igh and Cd83 bait in activated B cells and Tcrb (Eb) bait in double negative T cells and immature B cells. RNA-Seq and ATAC-Seq experiments in DN and Immature B cells.

4C测序(4C-Seq)已被证实为一项强大的技术,可在全基因组范围内鉴定与单个靶标位点(bait)的相互作用,此类互作对基因调控具有重要意义。 然而,由于该技术固有的多种偏倚,4C-Seq数据分析往往极具挑战性。处理4C-Seq数据时,一项关键考量是:基因组各区段的信号分辨率差异,源于其与靶标位点的三维空间距离差异——这会导致靶标位点紧邻区域的信号强度达到峰值,而远端顺式(far-cis)与反式(trans)区域的信号强度随距离增加逐渐降低。 4C-Seq实验的另一关键特征为分辨率,其分辨率水平极大程度上受限制性内切酶的选择及其在基因组中的切割频率影响。因此,理想的4C-Seq分析方法应具备足够的灵活性,可适配不同酶系处理产生的数据集,并能够完成全基因组范围的互作区域鉴定。 当前主流的4C-Seq分析方法仅能鉴定靶标位点附近区域或远端顺式、反式区域的互作,尚无方法能够全面解析不同长度尺度下的4C信号。此外,部分方法在使用高频限制性内切酶制备染色质片段的实验中同样无法正常运行。 本文介绍的4C-ker是一款基于隐马尔可夫模型(Hidden-Markov Model)的分析流程,可用于鉴定全基因组范围内与4C靶标位点发生互作的区域。同时,该流程整合了用于识别不同基因型或实验条件下获取的多组4C-Seq数据集间差异互作的分析方法,并通过自适应窗口大小校正靶标邻近区域、远端顺式及反式染色体上的信号覆盖度差异。 依托多组独立数据集验证,我们证实4C-ker在可重复鉴定不同分辨率酶系处理下全基因组范围互作结构域的性能上,优于所有现有4C-Seq分析流程。本研究使用的实验数据集包括:活化B细胞中Igh与Cd83靶标位点的4C-Seq实验,双阴性T细胞与未成熟B细胞中Tcrb (Eb)靶标位点的4C-Seq实验,以及双阴性T细胞和未成熟B细胞的RNA测序(RNA-Seq)与转座酶可及性测序(ATAC-Seq)实验。

二维码
社区交流群
二维码
科研交流群
商业服务