遇见数据集

Differential Gene Expression in the Siphonophore <i>Nanomia bijuga</i> (Cnidaria) Assessed with Multiple Next-Generation Sequencing Workflows

收藏
NIAID Data Ecosystem2026-03-07 收录
官方服务:

资源简介:

We investigated differential gene expression between functionally specialized feeding polyps and swimming medusae in the siphonophore Nanomia bijuga (Cnidaria) with a hybrid long-read/short-read sequencing strategy. We assembled a set of partial gene reference sequences from long-read data (Roche 454), and generated short-read sequences from replicated tissue samples that were mapped to the references to quantify expression. We collected and compared expression data with three short-read expression workflows that differ in sample preparation, sequencing technology, and mapping tools. These workflows were Illumina mRNA-Seq, which generates sequence reads from random locations along each transcript, and two tag-based approaches, SOLiD SAGE and Helicos DGE, which generate reads from particular tag sites. Differences in expression results across workflows were mostly due to the differential impact of missing data in the partial reference sequences. When all 454-derived gene reference sequences were considered, Illumina mRNA-Seq detected more than twice as many differentially expressed (DE) reference sequences as the tag-based workflows. This discrepancy was largely due to missing tag sites in the partial reference that led to false negatives in the tag-based workflows. When only the subset of reference sequences that unambiguously have tag sites was considered, we found broad congruence across workflows, and they all identified a similar set of DE sequences. Our results are promising in several regards for gene expression studies in non-model organisms. First, we demonstrate that a hybrid long-read/short-read sequencing strategy is an effective way to collect gene expression data when an annotated genome sequence is not available. Second, our replicated sampling indicates that expression profiles are highly consistent across field-collected animals in this case. Third, the impacts of partial reference sequences on the ability to detect DE can be mitigated through workflow choice and deeper reference sequencing.

本研究针对刺胞动物门(Cnidaria)管水母类(siphonophore)物种双小体管水母(Nanomia bijuga)中功能特化的摄食水螅体(feeding polyps)与游动钟形体(swimming medusae)之间的差异基因表达(differential gene expression)展开研究,采用长读长测序(long-read sequencing)/短读长测序(short-read sequencing)混合策略。本研究从长读长测序数据(罗氏454,Roche 454)中组装得到一套部分基因参考序列(gene reference sequences),并通过重复组织样本生成短读长序列,将其比对至参考序列以量化基因表达水平。本研究收集并对比了三种基于短读长的基因表达分析流程(short-read expression workflows)所得到的表达数据,这三种流程在样本制备、测序技术及比对工具上均存在差异:其一为Illumina mRNA-Seq,该方法沿每条转录本(transcript)的随机位置生成序列读段(sequence reads);另外两种为基于标签的分析方法(tag-based approaches)SOLiD SAGE与Helicos DGE,二者均从特定标签位点生成序列读段。不同流程间的表达分析结果差异,主要源于部分基因参考序列中数据缺失带来的差异化影响。当纳入所有由454测序得到的基因参考序列时,Illumina mRNA-Seq所检测到的差异表达(differentially expressed, DE)参考序列数量是基于标签的流程的两倍以上。这一差异主要源于部分参考序列中缺失标签位点,进而导致基于标签的流程出现假阴性(false negatives)结果。当仅纳入明确携带标签位点的参考序列子集时,三种流程间展现出高度一致性,且均识别出了相似的差异表达参考序列集合。本研究结果在多个方面为非模式生物(non-model organisms)的基因表达研究提供了可行参考:其一,本研究证明,在缺乏注释基因组序列(annotated genome sequence)的情况下,长读长测序/短读长测序混合策略是获取基因表达数据的有效手段;其二,本研究的重复采样实验表明,在该物种中,野外采集个体的表达谱(expression profiles)具有高度一致性;其三,通过选择合适的分析流程以及对参考序列进行更深层次测序,可缓解部分参考序列对差异表达基因检测能力带来的负面影响。

创建时间:
2016-10-28
二维码
社区交流群
二维码
科研交流群
商业服务