Supporting data for "Is there a fly in my soup? To what extent do individual and metabarcoding tell the same story?"
收藏资源简介:
This repository holds raw data and scripts for the analysis in Is there a fly in my soup? To what extent do individual and metabarcoding tell the same story?. Raw data for individual barcoding (originally published in A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level). To replicate analysis, all of these should be saved in a subdirectory "data" under the project directory. LIFEPLAN EPA Sequel data all reads - May 2022.xlsx - PacBio Sequel reads for all individuals, including reads classified as off-target. DS-LPEPA22.fas - Dominant sequence for each individual. Obtained from https://bench.boldsystems.org; Select dataset DS-LPEPA22; Downloads -> Sequences LP-BOLD_2304.xlsx - Metadata for each individual. Obtained from https://bench.boldsystems.org; Select dataset DS-LPEPA22; Downloads -> Data spreadsheets; Select all fields, format "multi-page" Script to download raw data for metabarcoding (originally published in A multi-modal dataset for insect biodiversity with imagery and DNA at the trap and individual level). These are not required to replicate the results unless the metabarcoding pipelines are to be re-run. assID45_ENA_accnos.tsv - Table of European Nucleotide Archive (ENA) accession numbers for each sample. download_MassID45_ENA.sh - Shell script to download raw FASTQ files from ENA. Additional input data and summary results for metabarcoding. These files should be saved in a subdirectory "data" under the project directory. LIFEPLAN_00001_UMIMap.txt and LIFEPLAN_00002_UMIMap.txt - Maps between primer indices and samples. These are used to correct some data/metadata mismatches. Analysis scripts. May require installation of R packages. 1_load_data.R - read raw data and produce combined.Rdata, which contains all input for subsequent scripts. This script cannot be directly reproduced, since some of the metabarcoding data is not publically available, since the sequencing runs included samples from unreleased studies. However it is included here for transparency. 2_dissimilarity.R - calculates dissimilarity matrices and saves results as dissimilarity.Rdata. 3_ref_presence_model.R - fits presence-absence models for the reference-based (MBRAVE) metabarcoding data and saves reults as ref_presence_model.Rdata. 4_ref_abundance_model.R - fits abundance models for the reference-based (MBRAVE) metabarcoding data and saves results as ref_abundance_model.Rdata. 5_denovo_presence_model.R - fits presence-absence models for the de novo (OptimOTU) metabarcoding data and saves results as denovo_presence_model.Rdata. 6_denovo_abundance_model.R - fits abundance models for the de novo (OptimOTU) metabarcoding data and saves results as denovo_abundance_model.Rdata. Result files generated by above analysis scripts. Should be placed in ubdirectory "data" under the project directory in order to generate summaries and figures. combined.Rdata - Contains the following R objects: study_bottles - character(45) - names of the 45 study bottles all_sample_rep_table - data.frame, 207 × 5 - all sample replicates all_sample_table - data.frame, 69 × 4 - all samples mbrave_rep_counts - data.frame, 37205 × 16 - read counts for each BIN in each sample from MBRAVE bold_data_full - data.frame, 36402 × 8 - individual arthropod sample data from BOLD bold_counts - data.frame, 7818 × 3 - counts of individuals for each BIN in eace sample from BOLD ref_rep_counts - data.frame, 40387 × 8 - combined MBRAVE + BOLD, per sample replicate ref_intersect_counts - data.frame, 10518 × 7 - combined MBRAVE + BOLD, per sample (rep intersection) ref_union_counts - data.frame, 17553 × 7 - combined MBRAVE + BOLD, per sample (rep union) ref_rep_richness - data.frame, 207 × 13 - combined MBRAVE + BOLD, BIN richness per sample replicate ref_intersect_richness - data.frame, 69 × 12 - combined MBRAVE + BOLD, BIN richness per sample (rep intersection) ref_union_richness - data.frame, 69 × 12 - combined MBRAVE + BOLD, BIN richness per sample (rep union) denovo_meta_rep_counts - data.frame, 16061 × 10 - read count each OTU in each sample replicate from OptimOTU denovo_ind_counts - data.frame, 7421 × 4 - individual arthropod counts per OTU from OptimOTU denovo_rep_counts - data.frame, 25940 × 8 - read counts for each OTU in each sample replicate from OptimOTU denovo_union_counts - data.frame, 9225 × 7 - combined OptimOTU, per sample (rep union) denovo_intersect_counts - data.frame, 8206 × 7 - combined OptimOTU, per sample (rep intersection) denovo_rep_richness - data.frame, 207 × 13 - OTU richness per sample replicate denovo_intersect_richness - data.frame, 69 × 12 - OTU richness per sample (rep intersection) denovo_union_richness - data.frame, 69 × 12 - OTU richness per sample (rep union) bin_order - data.frame, 6507 × 2 - mapping from BIN to taxonomic order (Arthropoda) bin_taxonomy - data.frame, 6565 × 9 - combined BIN taxonomy from MBRAVE + BOLD otu_order - data.frame, 19421 × 2 - mapping from OTU to taxonomic order (Arthropoda) otu_taxonomy - data.frame, 19442 × 10 - taxonomy for de novo OTUs dissimilarity.Rdata - Contains the following R objects: ref_meta_rep_dissim_matrix - data.frame, 6075 × 14 - dissimilarity between MBRAVE bulk replicates ref_ivm_dissim_matrix - data.frame, 6075 × 12 - dissimilarity between MBRAVE bulk and individual barcoding ref_ind_dissim - data.frame, 990 × 10 - pairwise dissimilarities between samples from individual barcoding ref_union_dissim - data.frame, 990 × 10 - pairwise dissimilarities between samples from union of MBRAVE reps ref_intersect_dissim - data.frame, 990 × 10 - pairwise dissimilarities between samples from intersection of MBRAVE reps denovo_meta_rep_dissim_matrix - data.frame, 6075 × 14 - dissimilarity between OptimOTU bulk replicates denovo_ivm_dissim_matrix - data.frame, 6075 × 12 - dissimilarity between OptimOTU bulk and individual barcoding denovo_ind_dissim - data.frame, 990 × 10 - pairwise dissimilarities between samples from individual barcoding denovo_union_dissim - data.frame, 990 × 10 - pairwise dissimilarities between samples from union of OptimOTU reps denovo_intersect_dissim - data.frame, 990 × 10 - pairwise dissimilarities between samples from intersection of OptimOTU reps ref_presence_model.Rdata - Contains the following R objects: ref_indpres_metaabund - data.frame, 15210 × 17 - Individual presence vs. MBRAVE metabarcoding abundances ref_indpres_m1 - glmmTMB - Model 1 -- fixed effects only ref_indpres_m1_predict - data.frame, 405 × 3 - Model 1 -- fixed effects only -- predictions ref_indpres_m2 - glmmTMB - Model 2 -- taxonomic random effect ref_indpres_m2_predict - data.frame, 4455 × 6 - Model 2 -- taxonomic random effect -- predictions ref_indpres_m2_loo - data.frame, 15210 × 17 - Model 2 -- taxonomic random effect -- leave-one-out predictions ref_indpres_m3 - glmmTMB - Model 3 -- taxonomic and bottle random effects ref_metapres_indabund - data.frame, 7818 × 16 - MBRAVE metabarcoding presence vs. individual abundances ref_metapres_m1 - glmmTMB - Model 1 -- fixed effects only ref_metapres_m1_predict - data.frame, 210 × 3 - Model 1 -- fixed effects only -- predictions ref_metapres_m2 - glmmTMB - Model 2 -- taxonomic random effect ref_metapres_m2_predict - data.frame, 1890 × 6 - Model 2 -- taxonomic random effect -- predictions ref_metapres_m2_loo - data.frame, 7818 × 16 - Model 2 -- taxonomic random effect -- leave-one-out predictions ref_metapres_m3 - glmmTMB - Model 3 -- taxonomic and bottle random effects ref_abundance_model.Rdata - Contains the following R objects: ref_hurdle_counts - data.frame, 15210 × 16 - MBRAVE individual abundances vs. metabarcoding reads ref_hurdle_m1 - glmmTMB - Model 1 -- fixed effects only ref_hurdle_m1_predict - data.frame, 405 × 3 - Model 1 -- fixed effects only -- predictions ref_hurdle_m2 - glmmTMB - Model 2 -- taxonomic random effect ref_hurdle_m2_predict - data.frame, 4455 × 6 - Model 2 -- taxonomic random effect -- predictions ref_hurdle_m2_loo - data.frame, 15210 × 17 - Model 2 -- taxonomic random effect -- leave-one-out predictions ref_hurdle_m3 - glmmTMB - Model 3 -- taxonomic and bottle random effects denovo_presence_model.Rdata - Contains the following R objects: denovo_indpres_metaabund - data.frame, 6409 × 18 - Individual presence vs. OptimOTU metabarcoding abundances denovo_indpres_m1 - glmmTMB - Model 1 -- fixed effects only denovo_indpres_m1_predict - data.frame, 381 × 3 - Model 1 -- fixed effects only -- predictions denovo_indpres_m2 - glmmTMB - Model 2 -- taxonomic random effect denovo_indpres_m2_predict - data.frame, 3048 × 6 - Model 2 -- taxonomic random effect -- predictions denovo_indpres_m2_loo - data.frame, 6409 × 18 - Model 2 -- taxonomic random effect -- leave-one-out predictions denovo_indpres_m3 - glmmTMB - Model 3 -- taxonomic and bottle random effects denovo_metapres_indabund - data.frame, 7377 × 17 - OptimOTU metabarcoding presence vs. individual abundances denovo_metapres_m1 - glmmTMB - Model 1 -- fixed effects only denovo_metapres_m1_predict - data.frame, 216 × 3 - Model 1 -- fixed effects only -- predictions denovo_metapres_m2 - glmmTMB - Model 2 -- taxonomic random effect denovo_metapres_m2_predict - data.frame, 1944 × 6 - Model 2 -- taxonomic random effect -- predictions denovo_metapres_m2_loo - data.frame, 7377 × 17 - Model 2 -- taxonomic random effect -- leave-one-out predictions denovo_metapres_m3 - glmmTMB - Model 3 -- taxonomic and bottle random effects denovo_abundance_model.Rdata - Contains the following R objects: denovo_hurdle_counts - data.frame, 6409 × 17 - OptimOTU individual abundances vs. metabarcoding reads denovo_hurdle_m1 - glmmTMB - Model 1 -- fixed effects only denovo_hurdle_m1_predict - data.frame, 381 × 3 - Model 1 -- fixed effects only -- predictions denovo_hurdle_m2 - glmmTMB - Model 2 -- taxonomic random effect denovo_hurdle_m2_predict - data.frame, 3048 × 6 - Model 2 -- taxonomic random effect -- predictions denovo_hurdle_m2_loo - data.frame, 6409 × 18 - Model 2 -- taxonomic random effect -- leave-one-out predictions denovo_hurdle_m3 - glmmTMB - Model 3 -- taxonomic and bottle random effects Summary of results and figure generation. is_there_a_fly_in_my_soup_analysis_fignums_revised.Rmd - R markdown document which generates summary information quoted in the text, as well as tables and figures. is_there_a_fly_in_my_soup_analysis_fignums_revised.pdf - Output of the above R markdown document.



