Additional file 1: of Comparative genomics of Bacteria commonly identified in the built environment
收藏资源简介:
Table S1. Selection of bacterial genera commonly identified in the built environment. Bacterial genera identified in 54 publications were compiled (see Table S2) and commonly identified genera were selected. All bacterial genera identified in more than about 10% of the publications (n ≥ 6 publications) with at least one complete reference genome on the NCBI RefSeq database were used in this study (n = 28 genera). Table S2. Metadata for each reference. 54 publications were compiled, including metadata for location, sub-locations, bacterial genera identified, sample type, climate (Table S4 and S5), temperature (°C), and humidity (%). If temperature or humidity was not described by the publication, the average over a certain period of time (either the timeframe stated in the publication or the publication year) was obtained from online sources. Table S3. Publication count for each “Common BE Bacterial Genus” by macro-Level BE location. Macro-level BE Locations included indoor, outdoor, underground, and extreme. Further division by type of sample is also depicted, including surface (S), air (A), water (W). Darker orange color indicates more references identified the genera in the macro BE location and sample type while lighter orange color indicates fewer references. The total number of references for each location and genera are also shown. Table S4. Köppen climate classification. Köppen climate classification was used to identify the climate for each publication’s study location. Only the climate assignment between 1981 and 2010 was used for this study. Abbreviation descriptions, latitude, and longitude values are listed. Table S5. Publication count for each “Common BE Bacterial Genus” by climate. The climate was identified for each publication’s study location based on the closest Köppen latitude and longitude values and correlated with the Köppen ID (see Table S4 for Köppen assignment). For publications describing general locations (e.g., only provided a U.S. state name), a central location in the region was chosen for latitude and longitude. Publications without location specifics were not included, and publications in space were separated out to “Space” category. Darker orange color indicates more references identified the genera in the macro BE location and sample type while lighter orange color indicates fewer references. The total number of references for each location and genera are also shown. Table S6. MetaMetaDB environmental category assignment for each “Common BE Bacterial Genus.” MetaMetaDB is a database to search for the possible habitats a microorganism could live in and was made by collecting 16S rRNA sequences. Environmental categories for each “Common BE bacterial genus” were based on the identity threshold of 97%, corresponding to the species taxonomic level. Every species for each “Common BE genus” is listed with the corresponding environmental category, where “Y” indicates that the species has been previously identified in the category and “N” indicates the species has not been identified in the category. “Hits” indicates the number of 16S rRNA sequences used by the database. Table S7. Mean distance (Dmean) between all pairs of bacterial species for each “Common BE Bacterial Genus.” The Dmean was used to describe the genetic diversity among species within a genus. The genetic distance between a pair of bacteria was calculated with the K80 model using the ‘dist.dna’ function of the ‘ape’ package of R ( https://cran.r-project.org/web/packages/ape ). We used a nucleotide sequence alignment of the 16S rRNA genes in ‘The All-Species Living Tree’ Project ( https://www.arb-silva.de/projects/living-tree /). LTP datasets based on SILVA release 128 were downloaded from Archive ( https://www.arb-silva.de/no_cache/download/archive/living_tree/LTP_release_128 /). Bacterial genera for which 3 or more taxa (N > 2) were available at LTP_release_128 were included in the 16S rRNA diversity analysis. Table S8. Genome information. Genome features reported include size (Mb), GC content (%), GCSI (GC skew index), and S value (strength of selected codon usage). A genus was deemed BE if observed in at least 6 publications out of 54. The column “BE” shows the number of references that identified the genera. Table S9. Robustness of the study. The genome data set used in this study was tested over two levels: 1) different subsets of bacteria (e.g., Phyla of Proteobacteria, Firmicutes, and Actinobacteria) and also randomly selecting one representative for species that have multiple strains sequenced, and 2) testing different numbers of publications (n = 1, 2, 3, 4, 5, and 6) to select for BE genera. Table S10. Genomic feature statistical analysis for each MetaMetaDB selected environmental category. Each genomic feature per MetaMetaDB environmental category was analyzed to determine statistical significance between the “Common BE genomes” associated with an environment and the “Common BE genomes” not associated. Significance is indicated by q-value < 0.05 and large effect size by Cliff’s delta |d| > 0.474. (XLSX 3660 kb)
表S1. 常见建成环境细菌属筛选。本研究汇总了54篇文献中鉴定出的细菌属(详见表S2),并筛选出其中的常见类群。本研究纳入了在至少约10%的文献(n≥6篇)中被鉴定出,且在NCBI参考序列数据库(NCBI RefSeq)上至少拥有一条完整参考基因组的细菌属,共计28个属。 表S2. 各参考文献元数据。本研究汇总了54篇文献的元数据,包括采样地点、亚地点、鉴定出的细菌属、样本类型、气候(详见表S4和表S5)、温度(℃)及湿度(%)。若文献未提及温度或湿度,则从在线数据源获取该时段(文献中注明的时段或文献发表年份)的平均温湿度数据。 表S3. 按宏观建成环境位置分类的“常见建成环境细菌属”文献计数。宏观建成环境位置包括室内、室外、地下及极端环境。同时还按样本类型进一步划分,包括表面(S)、空气(A)和水(W)。颜色越深代表在该宏观建成环境位置及样本类型中鉴定出该属的参考文献越多,颜色越浅则代表相关参考文献越少。此外还标注了各位置和属对应的参考文献总数。 表S4. 柯本气候分类法(Köppen climate classification)。本研究采用柯本气候分类法对各文献的研究地点进行气候归类,本研究仅使用1981-2010年间的气候分类结果。文中列出了分类缩写说明、纬度及经度数值。 表S5. 按气候分类的“常见建成环境细菌属”文献计数。本研究根据各文献研究地点的最近柯本气候分类纬度和经度值确定其气候类型,并与柯本气候编码相对应(柯本分类赋值详见表S4)。对于仅给出大致位置(例如仅提供美国州名)的文献,选取该区域的中心位置作为其经纬度。未提供具体位置信息的文献被排除,而太空相关研究的文献被单独归入“太空”类别。颜色越深代表在该宏观建成环境位置及样本类型中鉴定出该属的参考文献越多,颜色越浅则代表相关参考文献越少。此外还标注了各位置和属对应的参考文献总数。 表S6. “常见建成环境细菌属”的MetaMetaDB环境类别赋值。MetaMetaDB是一款用于检索微生物潜在生存生境的数据库,其构建基于16S rRNA基因(16S rRNA gene)序列的收集。本研究中“常见建成环境细菌属”的环境类别基于97%的序列相似度阈值确定,该阈值对应物种分类水平。本研究列出了每个“常见细菌属”的所有物种及其对应的环境类别,其中“Y”代表该物种此前已在该类别中被鉴定出,“N”代表未在该类别中被鉴定出。“Hits”代表数据库使用的16S rRNA序列数量。 表S7. 各“常见建成环境细菌属”所有细菌物种对之间的平均遗传距离(Dmean)。Dmean用于描述某一属内物种间的遗传多样性。细菌物种对之间的遗传距离采用R语言的ape包(ape package)中的`dist.dna`函数,基于K80模型计算得到。本研究使用了“全物种生命树”项目(The All-Species Living Tree Project)中的16S rRNA基因核苷酸序列比对数据,相关链接为https://www.arb-silva.de/projects/living-tree/。基于SILVA 128版的LTP数据集可从存档站点下载,链接为https://www.arb-silva.de/no_cache/download/archive/living_tree/LTP_release_128/。本研究仅纳入在LTP_release_128中拥有3个及以上分类单元(N>2)的细菌属用于16S rRNA多样性分析。 表S8. 基因组信息。本研究报告的基因组特征包括基因组大小(Mb)、GC含量(%)、GC偏倚指数(GC skew index)以及S值(密码子使用偏好强度)。若某一细菌属在54篇文献中至少被6篇文献检测到,则将其判定为建成环境相关细菌属。“BE”列显示了鉴定出该属的参考文献数量。 表S9. 本研究的稳健性检验。本研究使用的基因组数据集从两个层面进行了稳健性测试:1)选取不同的细菌类群子集(例如变形菌门、厚壁菌门和放线菌门),并对于存在多个测序菌株的物种随机选取一个代表菌株;2)选取不同数量的文献(n=1、2、3、4、5、6)用于筛选建成环境相关细菌属。 表S10. 各MetaMetaDB选定环境类别的基因组特征统计分析。本研究针对每个MetaMetaDB环境类别下的基因组特征进行分析,以确定与某一环境相关的“常见建成环境基因组”和不相关的“常见建成环境基因组”之间的统计学显著性差异。显著性以q值<0.05作为判定标准,效应量大小以克利夫斯Δ值(Cliff's delta)|d|>0.474作为大效应量的判定标准。(XLSX 3660 kb)



