遇见数据集

Data and codes for "Variable performances of commercial eDNA inventories challenge their use for surveying stream fish communities"

收藏
Zenodo2026-04-29 更新2026-05-26 收录
官方服务:

资源简介:

# Data and Code from: Variable performances of commercial eDNA inventories challenge their use for surveying stream fish communities ## Overview This repository contains 1- the raw data for the local fish community (species name, standard body length, and body weight), and 2- the bioinformatic pipelines, statistical scripts, and intermediate data files used to reanalyse raw environmental DNA (eDNA) sequencing data provided by three anonymized commercial companies (Company A, Company B, and Company D), and ## Repository Structure The repository is divided into one Excel file with raw data for the local fish community ('FishComData'), and three main directories, corresponding to the three anonymized companies (`Company_A`, `Company_B`, and `Company_D`). Within each company's directory, the workflow is split into two distinct sub-folders: Bioinformatics_Pipeline: Contains the bash scripts used to process the raw demultiplexed or assembled reads (FASTQ files) through the OBITools and BLASTn pipelines. The raw FASTQ files for each company are located here. Statistical_Analysis: Contains the R scripts used for downstream quality filtering and final count aggregation. This folder also holds the necessary intermediate outputs generated by the bioinformatic pipeline (e.g., filtered FASTA sequences, tabular read counts, and raw BLAST assignments). ## Important Notes on Data & Dependencies 1. Reference Databases To ensure the reproducibility of the BLAST assignments in the bash scripts, users must use the comprehensive 12S DNA reference library for the freshwater fish of French Guiana. The FASTA reference files (`reference_taxonomy_db...`) are not included in this repository. Before running the bioinformatic pipelines, you can download the reference database from Brosse, S. et al. (2026). Near complete 12S DNA reference library for the freshwater fish of French Guiana, northern Amazonian region. 2. Taxonomic Nomenclature Please note that the taxonomic names generated in the final output tables by the R scripts have not been manually updated to reflect the absolute latest taxonomic revisions. Consequently, there may be slight nomenclature differences between the raw outputs of these scripts and the finalized species lists presented in the published manuscript. The manuscript reflects the standardized taxonomy according to Le Bail et al. (2026). 3. Anonymization To maintain the integrity of the blind test evaluated in our study, the identities of the commercial eDNA providers have been strictly anonymized as "Company A", "Company B", and "Company D". ## Instructions for Use Step 1: Bioinformatic Processing Navigate to the `Bioinformatics_Pipeline` folder of the desired company. The shell scripts (`.sh`) are designed to run on a cluster. They require `Miniconda3`, `OBITools` (v1.2.11), `Cutadapt` (v4.3), and `NCBI_Blast+` (v2.10.0+). Step 2: Statistical Analysis Navigate to the `Statistical_Analysis` folder. The R scripts process the outputs from Step 1. They calculate alignment coverage, apply stringent quality filters, resolve taxonomic conflicts and aggregate read counts by sample and taxon. Required R packages: `tidyverse`, `Biostrings`, `openxlsx`.

提供机构:
INRAE-DECOD
创建时间:
2026-04-29
二维码
社区交流群
二维码
科研交流群
商业服务