遇见数据集

Global biogeography of N<sub>2</sub>-fixing microbes: <i> </i><i>nifH</i> amplicon database and analytics workflow

收藏
Figshare2025-02-04 更新2026-04-08 收录
官方服务:

资源简介:

This figshare entry has the <i>nifH</i> ASV database described in Morando, Magasin et al. 2025. In this work we compiled published data for marine nitrogenase (<i>nifH</i>) amplicons that were sequenced on Illumina platforms, and ran the datasets through our DADA2 <i>nifH</i><i> </i>workflow. Briefly, the workflow entailed: detecting ASVs with a DADA2 pipeline customized for <i>nifH</i><i>;</i> quality filtering and annotating ASVs using resources maintained by the Zehr and Turk-Kubo labs at UC Santa Cruz; and co-locating the samples with environmental data from the Simons Collaborative Marine Atlas Project (CMAP).The <i>nifH</i> ASV database in the 2025 publication is nifH_ASV_database_v2.tgz. An earlier version of the database (v1) is described in the preprint. Future versions will appear in figshare as new amplicon datasets become available.Software used to create the <i>nifH</i> ASV database:nifH-ASV-workflow in GitHub. This provides an overview of the workflow, a map of all included samples, and the scripts used to generate the <i>nifH</i> ASV database.nifH_amplicons_DADA2 in GitHub. This contains our DADA2 pipeline which processed samples by study. Then ASVs across all studies were processed together by the workflow. The pipeline can be used with amplicon data for genes other than <i>nifH</i> by specifying the primers and changing one parameter for how DADA2 error models are created.<br>

本figshare条目收录了Morando、Magasin等学者2025年研究中所报道的<i>nifH</i>扩增子序列变异(Amplicon Sequence Variant,ASV)数据库。 本研究整合了已公开的、基于Illumina平台测序的海洋固氮酶<i>nifH</i>扩增子数据,并将所有数据集导入我们定制的<i>nifH</i>专属DADA2分析流程进行处理。简言之,该分析流程包含以下环节:采用针对<i>nifH</i>定制的DADA2流程识别ASV;借助加州大学圣克鲁兹分校Zehr实验室与Turk-Kubo实验室维护的数据集资源,完成ASV的质量过滤与注释;将样本数据与西蒙斯海洋协作地图集项目(Simons Collaborative Marine Atlas Project,CMAP)提供的环境数据进行空间匹配。 本2025年研究对应的<i>nifH</i> ASV数据库文件为nifH_ASV_database_v2.tgz。该数据库的早期版本(v1)已在预印本中进行详细说明。后续当有新的扩增子数据集公开时,更新版本的数据库将发布于figshare平台。 构建该<i>nifH</i> ASV数据库所使用的软件资源如下: 1. GitHub仓库nifH-ASV-workflow:该仓库提供了本分析流程的整体概述、所有纳入样本的空间分布地图,以及用于生成<i>nifH</i> ASV数据库的脚本文件。 2. GitHub仓库nifH_amplicons_DADA2:该仓库包含我们的DADA2分析流程,可按研究项目逐个处理样本,随后由整体流程将所有研究的ASV数据进行整合分析。该流程支持通过指定引物序列、调整DADA2错误模型构建的单个参数,适配除<i>nifH</i>以外的其他基因的扩增子数据分析。

创建时间:
2024-10-18
二维码
社区交流群
二维码
科研交流群
商业服务