Aggregated frequencies of transcription initiations observed in FANTOM5 CAGE data on GRCm38, including alignments with low mapping qualities
收藏资源简介:
<strong>Overview</strong> Aligned reads of the FANTOM5 CAGE data have been used after filtering (ones with mapping quality less than 20 or percent identity less than 85% were discarded) for general purpose, resulting in the data set consisting of only the reads aligned with confidence. The filtering process made possible to interpret the data without ambiguity, however it also limited interpretation of paralogous or duplicated regions within the genome. Here all of the 5'-ends of the CAGE read alignments, including the ones with low mapping quality, were counted. The counts in the individual profiles were aggregated and summed up. This data set is produced for mouse, in the same way to the one for human data set http://doi.org/10.5281/zenodo.1410835 <strong>Data files</strong> The resulting data files are formatted as bigWig (https://genome.ucsc.edu/FAQ/FAQformat.html#format6.1). '*.fwd.bw' and '*.rev.bw' represent forward and reverse strand on the genome, respectively. <strong>Methods</strong> The BAM files under https://fantom.gsc.riken.jp/5/datafiles/reprocessed/mm10_v7/basic/ were subjected to 5'-end counting by bedtools v2.27.1 (https://github.com/arq5x/bedtools2), followed by conversion into bigWig with jksrc v366 (http://hgdownload.cse.ucsc.edu/admin/).
**概述** 本数据集采用经过滤处理的FANTOM5 CAGE比对reads(比对质量低于20或序列一致性低于85%的reads已被剔除)开展通用分析,最终仅保留具有可信比对结果的reads以构建数据集。该过滤流程虽可消除数据解读的歧义,但同时也限制了对基因组内旁系同源区域与重复区域的分析。本次研究对所有CAGE read比对序列的5'端进行计数,涵盖比对质量较低的序列,随后对各样本的计数谱进行汇总累加。本小鼠数据集的制备流程与人类数据集(http://doi.org/10.5281/zenodo.1410835)完全一致。 **数据文件** 最终生成的数据文件采用bigWig格式(https://genome.ucsc.edu/FAQ/FAQformat.html#format6.1)。其中,*.fwd.bw与*.rev.bw分别对应基因组的正链与负链。 **方法** 从https://fantom.gsc.riken.jp/5/datafiles/reprocessed/mm10_v7/basic/ 获取的BAM文件,经bedtools v2.27.1(https://github.com/arq5x/bedtools2)完成5'端计数后,再通过jksrc v366(http://hgdownload.cse.ucsc.edu/admin/)转换为bigWig格式。



