遇见数据集

Datasets for "Reference library readiness for eDNA metabarcoding: A multi-marker gap analysis and prioritization roadmap of Philippine marine fishes"

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

These datasets accompany the manuscript “Reference library readiness for eDNA metabarcoding: a multi-marker gap analysis and prioritization roadmap of Philippine marine fishes”. It contains the raw input files, repository data package, Python processing and retrieval scripts, intermediate harmonised tables, downstream R analysis scripts, and exported summary outputs required to reproduce and audit the taxonomic harmonisation, sequence-record retrieval, marker-coverage summaries, accession/record-depth analyses, family-level comparisons, provenance analyses, and manuscript figure generation presented in the study. The workflow begins with a curated species list used for WoRMS-based taxonomic harmonisation. That harmonisation step generates the standardised species backbone and related tracking files used to support downstream repository processing. The WoRMS-derived outputs, including the harmonised species table and repository retrieval name list, are then used as inputs for both the GenBank and BOLD workflows. The GenBank workflow uses the WoRMS-harmonised inputs to retrieve and screen repository records, producing a resolved record-level feature table and downstream summary products including per-species marker coverage, per-species accession depth, marker-depth summaries, marker-presence summaries, family-level coverage tables, marker-tier summaries, and diagnostic outputs. The BOLD workflow uses the same harmonised taxonomic inputs together with the archived public BOLD data package BOLD_Public.26-Sep-2025.tar.gz. This produces the initial BOLD record table and quality-control outputs. Because some target species were not represented in the public-package-derived record set, the workflow also incorporates a patching step based on a legacy minimal record file from a discontinued API retrieval workflow. The patched final record table is then used to generate resolved per-species marker coverage tables, per-species record-depth summaries, marker-depth summaries, marker-tier summaries, marker-presence summaries, family-level coverage tables, and diagnostic outputs. Downstream analyses are performed in R using the processed GenBank and BOLD outputs. These scripts generate marker-coverage summaries, “lack all four markers” summaries, accession/record-depth figures, family-level bar charts, family-level heatmaps, provenance summaries, provenance maps, and manuscript-ready figure exports. The R scripts therefore serve as both analysis and figure-generation scripts. A detailed inventory of all archived folders, files, and their roles in the workflow is provided in File_Index_and_Descriptions.pdf. The archive is organised to preserve the provenance of the manuscript from raw taxonomic input and repository-specific processing through to the intermediate tables, final analyses, and exported figures.

本数据集配套于论文《环境DNA(environmental DNA)参考库就绪性:菲律宾海洋鱼类多标记缺口分析与优先级规划路线图》(原英文标题:"Reference library readiness for eDNA: a multi-marker gap analysis and prioritization roadmap of Philippine marine fishes")。数据集包含可复现与审核本研究中分类学整合、序列记录检索、标记覆盖度总结、登录号/记录深度分析、科级比较、溯源分析以及论文图表生成所需的原始输入文件、序列数据库仓储数据包、Python处理与检索脚本、中间整合表格、下游R语言(R)分析脚本以及导出的汇总结果。 分析流程始于经整理的物种列表,该列表用于基于世界海洋物种登记册(World Register of Marine Species, WoRMS)的分类学整合操作。该整合步骤将生成标准化的物种骨干文件与相关跟踪文件,用于支撑后续的数据库数据处理流程。由世界海洋物种登记册生成的输出结果,包括整合后的物种表格与数据库检索名称列表,将作为基因银行(GenBank)与生命条形码数据系统(Barcode of Life Data Systems, BOLD)两类分析流程的输入数据。 基因银行分析流程使用经世界海洋物种登记册整合后的输入数据,检索并筛选数据库记录,生成解析后的记录级特征表格与下游汇总产物,包括物种特异性标记覆盖度、物种特异性登录号深度、标记深度总结、标记存在情况总结、科级覆盖度表格、标记层级总结以及诊断输出结果。 生命条形码数据系统分析流程使用相同的整合分类学输入数据,结合存档的公开生命条形码数据系统数据包BOLD_Public.26-Sep-2025.tar.gz开展分析,将生成初始的生命条形码数据系统记录表格与质量控制输出结果。由于部分目标物种未在公开数据包衍生的记录集中出现,该流程还纳入了基于已停用API检索流程遗留的最小记录文件的补全步骤。补全后的最终记录表格将用于生成解析后的物种特异性标记覆盖度表格、物种特异性记录深度总结、标记深度总结、标记层级总结、标记存在情况总结、科级覆盖度表格以及诊断输出结果。 下游分析采用R语言(R),基于处理后的基因银行与生命条形码数据系统输出结果开展。这些脚本将生成标记覆盖度总结、"缺失全部四类标记"总结、登录号/记录深度图表、科级柱状图、科级热图、溯源总结、溯源地图以及可直接用于论文的图表导出文件。因此,R语言脚本同时承担分析与图表生成的双重功能。 所有存档文件夹、文件及其在分析流程中的作用的详细清单详见File_Index_and_Descriptions.pdf文件。本存档的组织方式可保留论文的溯源信息,覆盖从原始分类学输入、数据库专属处理到中间表格、最终分析以及导出图表的全流程。

创建时间:
2026-08-10
二维码
社区交流群
二维码
科研交流群
商业服务