Data from: Sorting things out: assessing effects of unequal specimen biomass on DNA metabarcoding
收藏资源简介:
Environmental bulk samples often contain many different taxa that vary several orders of magnitude in biomass. This can be problematic in DNA metabarcoding and metagenomic high-throughput sequencing approaches, as large specimens contribute disproportionately high amounts of DNA template. Thus, a few specimens of high biomass will dominate the dataset, potentially leading to smaller specimens remaining undetected. Sorting of samples by specimen size (as a proxy for biomass) and balancing the amounts of tissue used per size fraction should improve detection rates, but this approach has not been systematically tested. Here, we explored the effects of size sorting on taxa detection using two freshwater macroinvertebrate bulk samples, collected from a low-mountain stream in Germany. Specimens were morphologically identified and sorted into three size classes (body size < 2.5 × 5, 5 × 10, and up to 10 × 20 mm). Tissue powder from each size category was extracted individually and pooled based on tissue weight to simulate samples that were not sorted by biomass (“Unsorted”). Additionally, size fractions were pooled so that each specimen contributed approximately equal amounts of biomass (“Sorted”). Mock samples were amplified using four different DNA metabarcoding primer sets targeting the Cytochrome c oxidase I (COI) gene. Sorting taxa by size and pooling them proportionately according to their abundance lead to a more equal amplification of taxa compared to the processing of complete samples without sorting. The sorted samples recovered 30% more taxa than the unsorted samples at the same sequencing depth. Our results imply that sequencing depth can be decreased approximately fivefold when sorting the samples into three size classes and pooling by specimen abundance. Even coarse size sorting can substantially improve taxa detection using DNA metabarcoding. While high-throughput sequencing will become more accessible and cheaper within the next years, sorting bulk samples by specimen biomass or size is a simple yet efficient method to reduce current sequencing costs.
环境混合样本通常包含诸多不同的分类群(taxa),其生物量存在数个数量级的差异。这在DNA宏条形码(DNA metabarcoding)和宏基因组高通量测序技术中会带来问题,因为大型标本会不成比例地贡献更多的DNA模板。因此,少量高生物量的标本将主导测序数据集,可能导致小型标本无法被检测到。通过标本尺寸(作为生物量的替代指标)对样本进行分选,并平衡不同尺寸组分的取样组织量,有望提升检测效率,但该方法尚未经过系统性验证。 本研究以采自德国一条低山地溪流的两份淡水大型无脊椎动物混合样本为材料,探究了尺寸分选对分类群检测的影响。研究人员通过形态学鉴定将标本分为三个尺寸等级:体长<2.5×5 mm、5×10 mm以及≤10×20 mm。对每个尺寸等级的组织粉末分别进行DNA提取,并基于组织重量进行混合,以此模拟未按生物量分选的样本(下称“未分选组”);此外,将各尺寸组分按每个标本贡献大致相等的生物量进行混合,构建“分选组”。本研究使用靶向细胞色素c氧化酶亚基I(Cytochrome c oxidase I, COI)基因的四套不同DNA宏条形码引物对上述模拟样本进行扩增。 与未分选的完整样本处理方式相比,按尺寸分选分类群并按其丰度比例混合,可使分类群的扩增更趋均衡。在相同测序深度下,分选组检测到的分类群数量比未分选组多30%。研究结果表明,若将样本分为三个尺寸等级并按标本丰度进行混合,测序深度可降低约五倍。即使是粗略的尺寸分选,也能显著提升DNA宏条形码技术的分类群检测效率。尽管未来几年高通量测序将愈发普及且成本更低,但按标本生物量或尺寸对混合样本进行分选,仍是一种简单高效的现有测序成本优化方案。



