Discovering Transcription Factor Binding Sites in Highly Repetitive Regions of Genomes with Multi-Read Analysis of ChIP-Seq Data
收藏资源简介:
Chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) is rapidly replacing chromatin immunoprecipitation combined with genome-wide tiling array analysis (ChIP-chip) as the preferred approach for mapping transcription-factor binding sites and chromatin modifications. The state of the art for analyzing ChIP-seq data relies on using only reads that map uniquely to a relevant reference genome (uni-reads). This can lead to the omission of up to 30% of alignable reads. We describe a general approach for utilizing reads that map to multiple locations on the reference genome (multi-reads). Our approach is based on allocating multi-reads as fractional counts using a weighted alignment scheme. Using human STAT1 and mouse GATA1 ChIP-seq datasets, we illustrate that incorporation of multi-reads significantly increases sequencing depths, leads to detection of novel peaks that are not otherwise identifiable with uni-reads, and improves detection of peaks in mappable regions. We investigate various genome-wide characteristics of peaks detected only by utilization of multi-reads via computational experiments. Overall, peaks from multi-read analysis have similar characteristics to peaks that are identified by uni-reads except that the majority of them reside in segmental duplications. We further validate a number of GATA1 multi-read only peaks by independent quantitative real-time ChIP analysis and identify novel target genes of GATA1. These computational and experimental results establish that multi-reads can be of critical importance for studying transcription factor binding in highly repetitive regions of genomes with ChIP-seq experiments.
染色质免疫共沉淀结合高通量测序(Chromatin immunoprecipitation followed by high-throughput sequencing, ChIP-seq)正快速取代染色质免疫共沉淀结合全基因组芯片平铺分析(chromatin immunoprecipitation combined with genome-wide tiling array analysis, ChIP-chip),成为绘制转录因子结合位点与染色质修饰图谱的首选方法。当前分析ChIP-seq数据的主流方法仅依赖于唯一比对至相关参考基因组的读段(即唯一比对读段,uni-reads),这可能导致高达30%的可比对读段被遗漏。本研究提出了一种通用方法,用于利用那些比对至参考基因组多个位置的读段(即多位置比对读段,multi-reads):该方法基于加权比对方案,将多位置比对读段分配为分数计数。通过使用人类STAT1与小鼠GATA1的ChIP-seq数据集,我们证实引入多位置比对读段可显著提升测序深度,检测到仅通过唯一比对读段无法识别的新型结合峰,并改善了可比对区域内结合峰的检测效果。我们通过计算实验探究了仅通过多位置比对读段检测得到的结合峰的各类全基因组特征。总体而言,仅由多位置比对读段分析获得的结合峰与唯一比对读段识别的结合峰具有相似的特征,仅其中绝大多数位于节段性重复区域内。我们进一步通过独立的实时定量染色质免疫共沉淀分析,验证了多个仅在GATA1多位置比对读段分析中被检测到的结合峰,并鉴定出GATA1的全新靶基因。上述计算与实验结果证实,在利用ChIP-seq实验研究基因组高度重复区域的转录因子结合时,多位置比对读段具有至关重要的意义。



