遇见数据集

Successful Recovery of Nuclear Protein-Coding Genes from Small Insects in Museums Using Illumina Sequencing

收藏
Figshare2016-01-15 更新2026-04-29 收录
官方服务:

资源简介:

In this paper we explore high-throughput Illumina sequencing of nuclear protein-coding, ribosomal, and mitochondrial genes in small, dried insects stored in natural history collections. We sequenced one tenebrionid beetle and 12 carabid beetles ranging in size from 3.7 to 9.7 mm in length that have been stored in various museums for 4 to 84 years. Although we chose a number of old, small specimens for which we expected low sequence recovery, we successfully recovered at least some low-copy nuclear protein-coding genes from all specimens. For example, in one 56-year-old beetle, 4.4 mm in length, our de novo assembly recovered about 63% of approximately 41,900 nucleotides in a target suite of 67 nuclear protein-coding gene fragments, and 70% using a reference-based assembly. Even in the least successfully sequenced carabid specimen, reference-based assembly yielded fragments that were at least 50% of the target length for 34 of 67 nuclear protein-coding gene fragments. Exploration of alternative references for reference-based assembly revealed few signs of bias created by the reference. For all specimens we recovered almost complete copies of ribosomal and mitochondrial genes. We verified the general accuracy of the sequences through comparisons with sequences obtained from PCR and Sanger sequencing, including of conspecific, fresh specimens, and through phylogenetic analysis that tested the placement of sequences in predicted regions. A few possible inaccuracies in the sequences were detected, but these rarely affected the phylogenetic placement of the samples. Although our sample sizes are low, an exploratory regression study suggests that the dominant factor in predicting success at recovering nuclear protein-coding genes is a high number of Illumina reads, with success at PCR of COI and killing by immersion in ethanol being secondary factors; in analyses of only high-read samples, the primary significant explanatory variable was body length, with small beetles being more successfully sequenced.

本研究针对馆藏于自然历史博物馆的小型干燥昆虫标本,对其核蛋白编码基因、核糖体基因及线粒体基因开展高通量Illumina测序(Illumina sequencing)分析。本次测序共纳入1头拟步甲科甲虫(Tenebrionid beetle)与12头步甲科甲虫(Carabid beetle),标本体长介于3.7至9.7 mm之间,各标本在不同博物馆的馆藏时长为4至84年不等。尽管我们选取了一批老旧小型标本,预期其序列回收率较低,但仍从所有标本中成功获取到至少部分低拷贝核蛋白编码基因。例如,在1头体长4.4 mm、馆藏56年的甲虫标本中,通过从头组装(de novo assembly)可获取目标67个核蛋白编码基因片段中约41900个核苷酸的63%,而采用参考基因组组装(reference-based assembly)则可获得70%的目标序列。即便在测序成功率最低的步甲科标本中,参考基因组组装仍可从67个核蛋白编码基因片段中获取34个至少达到目标长度50%的序列片段。对参考基因组组装所用替代参考序列的探索结果显示,参考序列引发的偏倚迹象极少。所有标本的核糖体基因与线粒体基因均几乎获取到完整序列。我们通过两种方式验证了序列的整体准确性:一是将测序所得序列与通过聚合酶链式反应(PCR)、桑格测序(Sanger sequencing)获取的序列(包括同种新鲜标本的序列)进行比对;二是通过系统发育分析检验序列在预测系统发育位置中的排布情况。虽检测到少量序列可能存在不准确之处,但此类问题极少影响标本的系统发育定位结果。尽管本次研究的样本量较小,但探索性回归分析结果显示,影响核蛋白编码基因回收率预测成功率的主导因素为Illumina测序读段(Illumina reads)数量较高;聚合酶链式反应扩增细胞色素C氧化酶亚基I(COI)的成功率、标本经乙醇浸泡处死的处理方式为次要影响因素。仅针对高读段样本开展分析时,体长则为具有显著影响的核心解释变量,体型较小的甲虫测序成功率更高。

创建时间:
2016-01-15
二维码
社区交流群
二维码
科研交流群
商业服务