Data and code from: Host-microbiome associations of native and invasive small mammals across a tropical urban-rural ecotone
收藏资源简介:
16Sv4 rRNA gene sequencing data Checklist for 16Sv4 rRNA gene sequencing data submitted to ENA for the 245 samples analysed in the manuscript - project PRJEB81284 This TSV file contains: the sample ID ("sample"), the study ID in ENA ("study"), information about the type of sequencing ("instrument_model", "library_name", "library_source", "library_selection", "library_strategy", "library_layout"), the file name of the forward read ("forward_file_name"), the md5 of the forward read ("forward _file_md5"), the file name of the reverse read ("reverse_file_name"), the md5 of the reverse read ("reverse _file_md5"). The sequencing data of 354 samples from various species of small mammals are stored in the project PRJEB81284, but only the 245 samples from Sundamys muelleri, Rattus rattus, Rattus norvegicus, Suncus murinus were analysed in the paper. FASTQ files of the paired-end reads (R1: forward read; R2: reverse read) can be downloaded File name: file_checklist_project_PRJEB81284.tsv Bioinformatics Bash script used to analyse the 16Sv4 sequences to ASVs (Amplicon Sequence Variants) using the QIIME 2 package This SH file contains the bash command lines to analyse the 16Sv4 sequences to ASVs using the QIIME 2 package. File name: script_qiime2_Borneo_sm_microbiota.sh Metadata required to run script_qiime2_Borneo_sm_microbiota This TSV file contains: the sample ID, the ID of the individual small mammal, the small mammal species, the town or village in which the individual was captured, the geographic coordinates in UTM system, the sex and age of the individual. File name: metadata_qiime2_Borneo_sm_microbiota Statistical analysis Data --> R_data_Borneo_sm_microbiota Individual information of small mammal faecal samples analysed in the study This CSV file contains data on the 245 small mammal individuals that were captured in Borneo between March 2012 and May 2013 and whose fecal content was sequenced for microbial analysis. The table contains 245 rows and 9 columns. The variables include: "Sample_ID": unique identifier for each faecal sample. "ID_individual": unique identifier for each individual (one Sample_ID corresponds to one ID_individual). "Species": small mammal species information, including genus and species. "District": district in which the individual was captured. "Town_village": town or village in which the individual was captured. "Coord_UTM_x": geographic coordinates in UTM system corresponding to the longitude of the capture location. "Coord_UTM_y": geographic coordinates in UTM system corresponding to the latitude of the capture location. "Sex": sex of the individual (female, male, unknown). "Age": stade of maturity of the individual (adult, immature, juvenile, subadult, unknown) After filtering in R ("data_preparation.R"), 236 samples remained. --> df_samples File name: dataset_samples.csv Trapping effort with information about environment of the trapping locations and presence-absence data for the four analysed small mammal species This CSV file contains 3541 rows, corresponding to unique trapping locations, and 20 columns describing the environment and providing information about the presence-absence of the four studied species. "coord_id": unique identifier for each trap location, formatted as longitude_latitude. "LC20m_xxx": columns describing the landcover types within 20m radii around the trapping locations: housing (compound and soil), sealed, soil, agriculture (grass and tree), garden (grass and tree), fallow (grass and tree), forest edge, forest, water, others. The columns contain numeric values ranging from 0 to 100, where 0 corresponds to no landcover type and 100 represents total landcover. "PA_smue": presence-absence data for Sundamys muelleri (0 = absence, 1 = presence). "PA_rr": presence-absence data for Rattus rattus (0 = absence, 1 = presence). "PA_rn": presence-absence data for Rattus norvegicus (0 = absence, 1 = presence). "PA_smur": presence-absence data for Suncus murinus (0 = absence, 1 = presence). "ID_individual": unique identifier for each individual. File name: dataset_environment.csv ASV abundance table for analysed smalla mammal faecal samples This text file contains the number of sequences of each bacterial ASV found in each analysed samples. The table contains 6109 rows corresponding to the ASVs and 247 columns corresponding to 1) unique code assigned to the ASV from QIIME 2 ("ASV_code"), 2-246) 245 analysed samples, 247) empty column that is removed during the filtering ("taxonomy"). After filtering in R ("data_preparation.R"), 1864 unique ASVs and 236 samples remained. --> asv_table File name: dataset_asv.txt ASV taxonomy table This TSV file contains the taxonomic classification of the found ASVs. Each ASVs has been affiliated to the SILVA database. The table contains 6109 rows corresponding to the ASVs found in the 245 analysed samples and 3 columns corresponding to 1) unique code assigned to the ASV from QIIME 2 ("ASV_code"), 2) concatenate taxonomic classification (from domain to species) ("taxon"), 3) confidence of the taxonomic classification ("confidence"). After filtering in R ("data_preparation.R"), 1864 unique ASVs remained. --> taxo_table File name: dataset_taxonomy.tsv ASV phylogenetic tree The phylogenetic tree of each ASV. File name: ASV_phylogenetic_tree.nwk Phylogenetic trees (from vertlife.org project) of all small mammal species captured in Borneo during the trapping effort made between March 2012 and May 2013 (Wells et al. 2014) These trees are used to create a final consensus tree. File name: host_species_trees.nex Colour palette for Figure 5 = ANCOMBC results Colour palette used in the plot displaying ANCOMBC results (Figure 5). File name: ANCOMBC_palette_families.csv Colour legend for bacterial families from the microbiome composition plot (Figure 2 and Figure S2) Legend extracted from the microbiome composition plot, used as legend for Figure 5. File name: legend_bacterial_families.rds Scripts --> R_scripts_Borneo_sm_microbiota The zip file contains various scripts corresponding to each step of the statistical analysis run in R. data_preparation.R: data filtering, preparation and inspection. relative_occurrence_probability.R: computation of relative occurrence probability for each host species using Generalised Additive Models (script to create Figure 1). density_function.R: function to calculate mode and 95% CI. legend_extraction_function.R: function to extract the legend from a plot, modification of the individual_legend function from the "microshades" package. microbiome_composition.R: analysis of microbiome composition at host species and individual level (script to create Figure 2). alpha_diversity.R: alpha diversity metrics calculation and Generalised Linear Models (GLM) testing the effect of Land Use Intensity (LUI) on alpha diversity metrics. computation_beta_diversity_metrics.R: beta diversity metrics computation. beta_diversity_GDMs.R: Generalised Dissimilarity Modelling testing the effect of host phylogenetic relatedness, LUI and spatial proximity on beta diversity metrics (script to create Figure 3). beta_diversity_NMDS_beta_dispersion_ANOSIM.R: beta diversity visualisation and analysis using beta dispersion, Analysis of Similarity, and GLM testing the effect of LUI on beta dispersion (script to create figure 4). ANCOM-BC.R: analysis to find differentially abundant ASVs in relation to LUI for each host species (script to create figure 5).
16Sv4 rRNA基因测序数据核查清单 针对本论文中分析的245份样本,提交至欧洲核苷酸档案库(European Nucleotide Archive, ENA)的16Sv4 rRNA基因测序数据相关核查清单——项目编号PRJEB81284 本制表符分隔值(TSV)文件包含以下内容:样本ID("sample")、ENA中的研究ID("study")、测序类型相关信息("instrument_model"、"library_name"、"library_source"、"library_selection"、"library_strategy"、"library_layout")、正向读取文件名("forward_file_name")、正向读取文件的MD5值("forward_file_md5")、反向读取文件名("reverse_file_name")、反向读取文件的MD5值("reverse_file_md5")。 项目PRJEB81284中存储了来自多种小型哺乳动物物种的354份测序数据,但本论文仅分析了其中4个物种的245份样本:苏门答腊巨鼠(*Sundamys muelleri*)、黑家鼠(*Rattus rattus*)、褐家鼠(*Rattus norvegicus*)、臭鼩(*Suncus murinus*)。可下载配对末端读取的FASTQ格式文件(R1为正向读取;R2为反向读取)。 文件名称:file_checklist_project_PRJEB81284.tsv ## 生物信息学分析脚本 使用QIIME 2包将16Sv4序列分析为扩增子序列变异体(Amplicon Sequence Variants, ASVs)的Bash脚本 本SH文件包含用于通过QIIME 2包将16Sv4序列转化为ASVs的Bash命令行代码。 文件名称:script_qiime2_Borneo_sm_microbiota.sh 运行上述脚本所需的元数据 本TSV文件包含以下内容:样本ID、小型哺乳动物个体ID、小型哺乳动物物种、捕获个体所在的城镇/村落、通用横轴墨卡托(Universal Transverse Mercator, UTM)坐标系下的地理坐标、个体性别与年龄。 文件名称:metadata_qiime2_Borneo_sm_microbiota ## 统计分析数据 ### 研究样本基础信息 **R_data_Borneo_sm_microbiota** 本研究分析的小型哺乳动物粪便样本的个体信息 本逗号分隔值(CSV)文件包含2012年3月至2013年5月间在婆罗洲捕获的245只小型哺乳动物个体的数据,这些个体的粪便内容物已进行微生物组测序分析。该表格包含245行与9列,变量包括: "Sample_ID":每份粪便样本的唯一标识符。 "ID_individual":每个个体的唯一标识符(一份样本对应一个个体ID)。 "Species":小型哺乳动物物种信息,包含属与种。 "District":个体捕获所在的行政区。 "Town_village":个体捕获所在的城镇/村落。 "Coord_UTM_x":捕获位置经度对应的UTM坐标系地理坐标。 "Coord_UTM_y":捕获位置纬度对应的UTM坐标系地理坐标。 "Sex":个体性别(雌性、雄性、未知)。 "Age":个体成熟阶段(成体、未成熟、幼体、亚成体、未知)。 经R脚本`data_preparation.R`过滤后,剩余236份样本,得到数据集`df_samples`。 文件名称:dataset_samples.csv ### 捕获努力量与环境数据 包含捕获位点环境信息及4个分析物种存在-缺失数据的捕获努力数据集 本CSV文件包含3541行(对应唯一捕获位点)与20列,用于描述捕获位点的环境信息,并提供4个研究物种的存在-缺失数据。 变量说明: "coord_id":每个捕获位点的唯一标识符,格式为“经度_纬度”。 "LC20m_xxx":描述捕获位点周围20米半径范围内土地覆盖类型的列,包括:住房(建筑群与土壤)、硬化地表、裸土、农业用地(草本与树木)、花园(草本与树木)、休耕地(草本与树木)、林缘、森林、水体、其他。列中数值范围为0至100,0代表无该土地覆盖类型,100代表完全被该类型覆盖。 "PA_smue":苏门答腊巨鼠(*Sundamys muelleri*)的存在-缺失数据(0=未检出,1=检出)。 "PA_rr":黑家鼠(*Rattus rattus*)的存在-缺失数据(0=未检出,1=检出)。 "PA_rn":褐家鼠(*Rattus norvegicus*)的存在-缺失数据(0=未检出,1=检出)。 "PA_smur":臭鼩(*Suncus murinus*)的存在-缺失数据(0=未检出,1=检出)。 "ID_individual":每个个体的唯一标识符。 文件名称:dataset_environment.csv ### ASV丰度表 用于分析的小型哺乳动物粪便样本的ASV丰度表 本文本文件包含每份分析样本中检出的各细菌ASV的序列计数。该表格包含6109行(对应ASVs)与247列,分别为:1) QIIME 2为ASV分配的唯一代码("ASV_code");2~246列:245份分析样本;247列:过滤过程中需移除的空列("taxonomy")。 经R脚本`data_preparation.R`过滤后,剩余1864个唯一ASVs与236份样本,得到数据集`asv_table`。 文件名称:dataset_asv.txt ### ASV分类学表 本TSV文件包含检出的ASVs的分类学注释信息。所有ASV均已比对至SILVA数据库。该表格包含6109行(对应245份分析样本中检出的ASVs)与3列,分别为:1) QIIME 2为ASV分配的唯一代码("ASV_code");2) 拼接的分类学注释信息(从界到种)("taxon");3) 分类学注释的置信度值("confidence")。 经R脚本`data_preparation.R`过滤后,剩余1864个唯一ASVs,得到数据集`taxo_table`。 文件名称:dataset_taxonomy.tsv ### ASV系统发育树 各ASV的系统发育树。 文件名称:ASV_phylogenetic_tree.nwk ### 宿主物种系统发育树 2012年3月至2013年5月间在婆罗洲捕获的所有小型哺乳动物物种的系统发育树(数据来自vertlife.org项目,Wells et al. 2014) 这些树用于构建最终的共识系统发育树。 文件名称:host_species_trees.nex ### 可视化配色与图例 用于图5(ANCOMBC分析结果)的配色方案 用于绘制ANCOMBC分析结果图(图5)的配色板。 文件名称:ANCOMBC_palette_families.csv 微生物组组成图(图2与图S2)的细菌科分类群图例 从微生物组组成图中提取的图例,用于图5的绘图。 文件名称:legend_bacterial_families.rds ### R统计分析脚本 **R_scripts_Borneo_sm_microbiota** 该压缩包包含本研究中R语言统计分析各步骤对应的多个脚本: - `data_preparation.R`:数据过滤、整理与可视化检查。 - `relative_occurrence_probability.R`:使用广义加性模型(Generalized Additive Models, GAM)计算各宿主物种的相对出现概率(用于生成图1的脚本)。 - `density_function.R`:用于计算众数与95%置信区间的函数。 - `legend_extraction_function.R`:从绘图中提取图例的函数,改编自"microshades"包中的`individual_legend`函数。 - `microbiome_composition.R`:宿主物种与个体水平的微生物组组成分析(用于生成图2的脚本)。 - `alpha_diversity.R`:α多样性指数计算,以及使用广义线性模型(Generalized Linear Model, GLM)检验土地利用强度(Land Use Intensity, LUI)对α多样性指数的影响。 - `computation_beta_diversity_metrics.R`:β多样性指数计算。 - `beta_diversity_GDMs.R`:广义差异建模(Generalized Dissimilarity Modelling, GDM),检验宿主系统发育亲缘关系、LUI与空间邻近性对β多样性指数的影响(用于生成图3的脚本)。 - `beta_diversity_NMDS_beta_dispersion_ANOSIM.R`:使用β离散度、相似性分析(Analysis of Similarity, ANOSIM)进行β多样性可视化与分析,以及使用GLM检验LUI对β离散度的影响(用于生成图4的脚本)。 - `ANCOM-BC.R`:针对各宿主物种,分析与LUI相关的差异丰度ASVs(用于生成图5的脚本)。



