DADA2 formatted 16S rRNA gene sequences for both bacteria & archaea
收藏资源简介:
<strong><em>This version is to stay up to date with the improvements and increase in16S rRNA gene sequences added to the GTDB release 207. Please read this post for the stats on the updates. </em></strong><strong><em>https://gtdb.ecogenomic.org/stats/r207</em></strong><strong><em>.</em></strong><strong><em> </em></strong> <strong><em>There has been no change to the RDP-RefSeq reference database</em></strong> <strong><em>If anyone has concerns with MAG extracted 16S rRNA gene contamination concerns, then I suggest that they contact the curators of GTDB themselves because it is outside of my role with these resources designed for DADA2 usage only. </em></strong> <strong><em>Another concern that was raised was the orientation of the DB sequences, to get past this problem please use the tryRC = TRUE argument in the assignTaxonomy command within DADA2, this will search your ASVs in the reverse complement as well. </em></strong> This Version was primarily updated because we have recently updated the RefSeq+RDP database and also included mitochondrial and eukaryotic 16S rRNA sequences. Also because I decided to include the required formats to be able to use the addSpecies command in DADA2. This command searches the database at 100% identity and has the flexibility to either get the best hit or multiple hits to your amplicon. I recommend it if you are using a single or 2 region amplicons of the 16S rRNA gene. These two combined bacterial and archaeal 16S rRNA gene sequence databases were collated from various sources and formatted for the purpose of using the "assignTaxonomy" command within the DADA2 pipeline. The data was converted to suite DADA2 format by Alishum Ali. RefSeq+RDP: This database contains 22433 bacterial, 1055 archaea and 99 eukaryotic full lengths16S rRNA gene sequences. It was compiled by <strong>Paul Greenfield </strong>on the <strong>06/11/2020</strong> from predominantly the NCBI RefSeq 16S rRNA database (https://www.ncbi.nlm.nih.gov/refseq/targetedloci/16S_process/) and was supplemented with extra sequences from the RDP database (https://rdp.cme.msu.edu/misc/resources.jsp). Genome Taxonomy Database (GTDB): The new version of our dada2 formatted GTDB reference sequences now contains 31319 bacteria and 1565 archaea full 16S rRNA gene sequences. If you wonder why there are fewer species with 16S rRNA, that is because some metagenomics assembled genomes (MAGs) lack the 16S gene and thus cannot be extracted. The database was downloaded from https://data.ace.uq.edu.au/public/gtdb/data/releases/ on 28/04/2020. Please read the release notes and file descriptions. The formatting to DADA2 was done using simple awk bash scripts. The script takes as input a fasta file and a tab-delimited taxonomy file (slightly edited to remove special characters) and then it outputs a fasta file with all 7 taxonomy ranks separated by ";" as required for DADA2 compatibility. Additionally, we have concatenated the unique sequence ID be it NCBI/RDP or GTDB ID to the species entry (but replaced the "." with an " _". We see this as an important QC step to highlight the issues/confidence associated with short read taxonomy assignment at the finer rank levels. Also, this update includes two other files that you can use with the assignTaxonomy and addSpecies commands in DADA2.
本版本旨在跟进基因组分类数据库(Genome Taxonomy Database, GTDB)207版的更新内容及新增的16S rRNA基因序列。有关本次更新的统计详情,请查阅该帖子:https://gtdb.ecogenomic.org/stats/r207。RDP-RefSeq参考数据库未作任何修改。若您对宏基因组组装基因组(Metagenome-Assembled Genomes, MAGs)提取的16S rRNA基因存在污染相关疑虑,建议直接联系GTDB的管理者——本类资源仅面向DADA2使用场景,相关问题超出我的职责范围。另有用户提出数据库序列方向相关的疑虑,解决该问题可在DADA2的"assignTaxonomy"命令中设置tryRC = TRUE参数,该参数将同时以反向互补序列检索您的扩增子序列变异(Amplicon Sequence Variants, ASVs)。本次版本更新主要出于两点原因:其一,我们近期更新了RefSeq+RDP数据库,并新增了线粒体及真核生物16S rRNA基因序列;其二,我补充了适配DADA2中"addSpecies"命令的必要文件格式。该命令将以100%相似度检索数据库,支持返回目的扩增子的最佳匹配或多个匹配结果。若您使用16S rRNA基因的单区域或双区域扩增子,推荐使用此命令。这两款整合了细菌与古菌16S rRNA基因序列的数据库,均从多源数据整理而来,并针对DADA2流程中的"assignTaxonomy"命令做了格式适配,该格式转换工作由Alishum Ali完成。RefSeq+RDP数据库:该数据库包含22433条细菌、1055条古菌及99条真核生物全长16S rRNA基因序列,由Paul Greenfield于2020年11月6日主要基于NCBI RefSeq 16S rRNA数据库(https://www.ncbi.nlm.nih.gov/refseq/targetedloci/16S_process/)整理,并补充了RDP数据库(https://rdp.cme.msu.edu/misc/resources.jsp)的额外序列。基因组分类数据库(GTDB):本次更新的DADA2格式GTDB参考序列库,现包含31319条细菌与1565条古菌全长16S rRNA基因序列。若您疑惑为何携带16S rRNA基因的物种数量较少,原因在于部分宏基因组组装基因组(MAGs)缺失16S基因,因此无法被提取。该数据库于2020年4月28日从https://data.ace.uq.edu.au/public/gtdb/data/releases/ 下载,请查阅其发布说明与文件描述。针对DADA2的格式转换工作通过简单的awk与bash脚本完成。该脚本以fasta文件与制表符分隔的分类学文件(经轻微编辑以移除特殊字符)作为输入,随后输出符合DADA2兼容性要求的fasta文件,其中7个分类学等级以分号(;)分隔。此外,我们将序列唯一标识符(无论是NCBI、RDP还是GTDB的ID)拼接至物种条目后(其中点号"."已替换为下划线"_")。我们认为这是一项重要的质量控制步骤,可用于凸显短读长序列在精细分类等级上的分类分配结果所存在的问题与置信度。此外,本次更新还新增了另外两款可用于DADA2中"assignTaxonomy"与"addSpecies"命令的文件。



