遇见数据集

Using soil eDNA for plant diversity assessments

收藏
Zenodo2022-04-01 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This data sets corresponds to a publication in Methods in Ecology and Evolution titled: <strong>Plant biodiversity assessment through soil eDNA reflects temporal and local diversity</strong> In August 2018, a single soil eDNA sample was collected from the centre of each permanent plot (1m2) in the Solhomfjell Forest Reserve stablished by the Sommerfeltia program. The soil eDNA samples were stored in individual plastic bags for transportation to the lab and stored at -20 °C prior to freeze-drying under vacuum. Each soil eDNA sample was separately homogenized with ceramic beads and one gram was used for eDNA extraction. The latter was done in five rounds of two steps: (1) CTAB/chloroform pre-treatment to increase the separation of the organic phase and (2) aqueous phase and using the E.Z.N.A. soil DNA kit following the manufacturer’s protocol (Omega Bio-tek, Norcross, Georgia, USA). The chloroplast marker trnL (UAA) intron P6 loop was chosen as its short sequence can yield amplification of old DNA material degraded in eDNA samples. This marker was amplified for each sample with the g and h primers by PCR, using three technical replicates (Taberlet et al. 2007; 5'-GGGCAATCCTGAGCCAA-3', 5'-CCATTGAGTCTCTGCACCTATC-3'). Forward and reverse primers were tagged with a unique 12 bp oligonucleotide on the 5’ end (Fadrosh et al. 2014). Unique combinations of tagged primers were set up in panels for each PCR reaction for a total of 309 samples (100 samples with 3 PCR replicates each, 5 extractions blanks and 4 PCR negatives). The PCR negatives had no DNA template and were placed on the 96th well position in each panel. Composition of PCR reactions, final volumes and number of cycles can be found in Supporting Information Data S1. The PCR products were run on a 2% agarose gel, and the amplicon concentrations were measured via band intensity using ImageLab software (Bio-Rad, California, USA). The lowest concentration (μM) available for all PCR products and its relative volume was identified and the relative concentrations of the PCR products were adjusted to this same concentration. Amplicons were pooled in one library using a Biomek 4000 automated liquid handler (Beckman Coulter Life Sciences, Indianapolis, Indiana, USA). The library was cleaned using AMPure XP reagent beads (Beckman Coulter Life Sciences, Indianapolis, Indiana, USA). The length for all amplicons in the library was determined using a Fragment Analyzer (Agilent Technologies, Santa Clara, California, USA). The library was sequenced on an Illumina MiSeq platform with 150 bp paired-end reads (Illumina Inc., San Diego, California, USA). Sequence data was analyzed and curated using OBITools 2 (Boyer et al. 2016) following the wolf tutorial with adaptations for demultiplexing dual indexes from QIIME2 (Caporaso et al. 2010). Sequences were retained with both indexes for dereplication for further analysis. Similar sequences were clustered with obiclean (Boyer et al. 2016) only when the read count of the less abundant sequence was below 5% of the most abundant sequence. To reduce multiple identifications of the same sequence, taxonomic assignment of dereplicated and denoised sequences was done by matching to three reference sequences databases containing: (i) only taxa registered in the local Solholmfjell reference library; (ii) the complete arctic boreal database for vascular plants and bryophytes (Sønstebø et al.2010; Willerslev et al. 2014; Soininen et al. 2015); and (iii) taxa available in the EMBL database (downloaded on 7/02/2020) filtered to sequences with trnL (UUA) intron g-h primers using ecoPCR tool from OBITools (Boyer et al. 2016). Resulting identifications from the three databases were merged by sequence and duplicates were eliminated giving priority to reference databases (i), (ii), and (iii) in that order. To minimize erroneous taxonomic assignments, only taxa with a 100% match to a reference sequence were retained. We observed that below this threshold, sequences remained without a taxonomic rank assigned. Further, assigned taxa names were changed to the lowest taxonomic rank possible with trnL (UUA) intron and thus are identical to those registered in vegetation surveys. When different sequences were identified with identical taxa names, a unique entry was retained and the read counts within plots and replicates were summed. Read counts were averaged across all samples and negative controls (extraction + PCR). All analyses are plot-based, and coded using R v 1.4.17 (R Core Team, 2019) and with packages listed in the code. Separate analyses are made for vascular plants and bryophytes, and/or for spruce and pine data subsets, or combinations thereof, when relevant.

本数据集对应发表于《Methods in Ecology and Evolution》(《生态学与进化方法》)的论文:《通过土壤环境DNA(environmental DNA, eDNA)开展植物多样性评估可反映局地与时间尺度的物种多样性》。2018年8月,在Sommerfeltia项目设立的索尔霍姆耶尔森林保护区内,于每个1平方米永久样地的中心采集了一份土壤eDNA样本。土壤eDNA样本分装于独立塑料袋中运输至实验室,并在真空冷冻干燥前于-20℃条件下保存。每份土壤eDNA样本均使用陶瓷珠单独均质化,称取1克用于eDNA提取。提取流程共开展五轮,每轮包含两个步骤:(1) 十六烷基三甲基溴化铵(CTAB)/氯仿前处理以增强有机相与水相的分离效果;(2) 采用E.Z.N.A.土壤DNA试剂盒(Omega Bio-tek,美国佐治亚州诺克罗斯市)按照制造商说明书完成后续操作。本研究选择叶绿体标记trnL (UAA)内含子P6环,因其序列较短,可对eDNA样本中降解的陈旧DNA物质进行扩增。使用g和h引物对每份样本的该标记进行聚合酶链式反应(PCR)扩增,设置3次技术重复(Taberlet等,2007;引物序列:5'-GGGCAATCCTGAGCCAA-3',5'-CCATTGAGTCTCTGCACCTATC-3')。正向与反向引物的5'端均连接有唯一的12 bp寡核苷酸标签(Fadrosh等,2014)。本次实验共设置309份样本的带标签引物唯一组合,每板对应一次PCR反应:其中100份样本各设置3次PCR重复,5份提取空白对照,4份PCR阴性对照。PCR阴性对照未添加DNA模板,置于每板的第96孔位。PCR反应的组分、终体积及循环次数详见补充信息数据S1。PCR产物经2%琼脂糖凝胶电泳分离,使用ImageLab软件(Bio-Rad,美国加利福尼亚州)通过条带强度测定扩增子浓度。确定所有PCR产物的最低浓度(μM)及其对应体积,将所有PCR产物的浓度统一调整至该水平。使用Biomek 4000自动化液体处理工作站(Beckman Coulter Life Sciences,美国印第安纳州印第安纳波利斯市)将扩增子混合构建测序文库。使用AMPure XP磁珠(Beckman Coulter Life Sciences,美国印第安纳州印第安纳波利斯市)对文库进行纯化。使用Fragment Analyzer(Agilent Technologies,美国加利福尼亚州圣克拉拉市)测定文库中所有扩增子的长度。随后在Illumina MiSeq平台(Illumina Inc.,美国加利福尼亚州圣迭戈市)上进行150 bp双端测序。序列数据分析与整理使用OBITools 2(Boyer等,2016)完成,参考wolf教程并针对QIIME2(Caporaso等,2010)的双索引解复用流程进行适配调整。保留同时包含双索引的序列进行去重复,用于后续分析。使用obiclean工具(Boyer等,2016)对相似序列进行聚类,仅当丰度较低序列的序列读长计数低于高丰度序列的5%时执行该聚类操作。为减少同一序列的多次鉴定,对去重复并降噪后的序列进行分类学注释时,将其与三个参考序列数据库进行比对:(i) 仅包含索尔霍姆耶尔本地参考文库中已登记的类群;(ii) 涵盖维管植物与苔藓植物的完整北极-北方区域数据库(Sønstebø等,2010;Willerslev等,2014;Soininen等,2015);(iii) EMBL数据库中于2020年2月7日下载的、通过OBITools的ecoPCR工具(Boyer等,2016)筛选出的包含trnL (UUA)内含子g-h引物结合位点的类群序列。将三个数据库的注释结果按序列合并,去除重复注释,优先顺序依次为数据库(i)、(ii)、(iii)。为降低错误的分类学注释结果,仅保留与参考序列完全匹配(100%一致)的类群。研究发现,低于该一致性阈值的序列无法被分配至明确的分类阶元。此外,将注释得到的类群名称调整至可通过trnL (UUA)内含子确定的最低分类阶元,其结果与植被调查中登记的名称一致。当不同序列被注释为相同类群名称时,仅保留唯一条目,并将对应样地与重复中的序列读长计数进行合并。对所有样本及阴性对照(提取空白+PCR阴性对照)的序列读长计数取平均值。所有分析均以样地为单位开展,使用R v 1.4.17(R核心团队,2019)及代码中列出的相关包完成。针对维管植物与苔藓植物、或云杉与松树数据子集,或相关组合分别开展独立分析。

提供机构:
Zenodo
创建时间:
2022-04-01
二维码
社区交流群
二维码
科研交流群
商业服务