遇见数据集

MUSIC

收藏
Figshare2015-12-14 更新2026-04-29 收录
官方服务:

资源简介:

High-throughput DNA sequencers are becoming indispensable in our understanding of diseases atmolecular level, in marker-assisted selection in agricultureand in microbial genetics research. These sequencinginstruments produce enormous amount of data (oftenterabytes of raw data in a month) that requires efficientanalysis, management and interpretation. The commonlyused sequencing instrument today produces billions ofshort reads (upto 150 bases) from each run. The first stepin the data analysis step is alignment of these short reads tothe reference genome of choice. There are different opensource algorithms available for sequence alignment to thereference genome. These tools normally have a highcomputational overhead, both in terms of number ofprocessors and memory. Here, we propose a hybrid computing environment called MUSIC (Mapping USInghybrid Computing) for one of the most popular open sourcesequence alignment algorithm, BWA, using acceleratorsthat show significant improvement in speed over the serialcode

高通量DNA测序仪在从分子层面解析疾病、开展农业标记辅助选择以及推进微生物遗传学研究等领域正成为不可或缺的核心工具。这类测序仪器可产生海量原始数据——单月原始数据量常达数十TB级别,亟需高效的分析、管理与解读方案。当前主流的测序平台每次运行可生成数十亿条短读段(最长可达150个碱基)。数据分析的首要步骤是将这些短读段比对至选定的参考基因组。目前已有多款开源算法可实现参考基因组序列比对,但这类工具通常存在较高的计算开销,无论是处理器核心需求还是内存占用均较为突出。针对当前最流行的开源序列比对算法之一BWA,本文提出一种名为MUSIC(Mapping USIng hybrid Computing,混合计算比对工具)的混合计算环境,借助加速硬件可实现相较于串行代码的显著性能提升。

创建时间:
2015-12-14
二维码
社区交流群
二维码
科研交流群
商业服务