Genome assemblies and respective wg/cgMLST profiles of a diverse dataset comprising 1,434 Salmonella enterica isolates
收藏资源简介:
<strong>Dataset</strong> This dataset comprises the genome assemblies and respective 8,558-loci whole-genome (wg) Multiple Locus Sequence Type (MLST) profiles [INNUENDO schema (Llarena et al. 2018) available in chewie-NS (Mamede et al. 2022)] of a final set of 1,434 <em>Salmonella enterica </em>samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA) at the beginning of the analysis (November 2021). This set of samples was carefully selected to cover a wide genetic diversity (assessed in terms of serotype). In total, 125 different serotypes are represented in this dataset, with Typhimurium (including monophasic), Enteritidis and Infantis being the most represented ones and, together, corresponding to 56.2% of the dataset. File “Se_metadata.xlsx” contains metadata information for each isolate, including ENA/SRA accession number, BioProject and in-silico MLST ST and serotype. The directory “assemblies/” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. The file “profiles/Se_profiles_wgMLST.tsv” corresponds to a tab separated file with the 8,558-loci wgMLST profiles of each isolate presented in the metadata file. The files “profiles/Se_profiles_cgMLST_95.tsv”, “profiles/Se_profiles_cgMLST_98.tsv” and “profiles/Se_profiles_cgMLST_100.tsv” correspond to a 3,261-loci, 3,179-loci and 874-loci cgMLST profiles of each isolate presented in the metadata file, respectively. These profiles were determined as explained below. <strong>Dataset selection and curation</strong> With the objective of creating a diverse dataset of <em>S. enterica</em> genome assemblies, we collected information about the genetic diversity (serotype) of the isolates available at Enterobase database in the beginning of this analysis (November 2021) and in other previous works. Based on this information, we selected an initial dataset comprising 1,779 samples associated with four BioProjects (PRJEB16326, PRJEB20997, PRJEB30335 and PRJEB39988). Their WGS data was downloaded from ENA/SRA with fastq-dl v1.0.6. Read quality control, trimming and assembly were performed with the Aquamis v1.3.9 (Deneke et al. 2021) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,434 isolates passed this curation step and were included in the final dataset. In-silico serotyping was performed with SeqSero2 v1.2.1 (Zhang et al. 2019). wgMLST profiles of each of these isolates were determined with chewBBACA v2.8.5 (Silva et al. 2018), using the 8,558-loci INNUENDO schema available in chewie-NS (Llarena et al. 2018; Mamede et al. 2022) and downloaded on May 31<sup>st</sup>, 2022. Three cgMLST schemas were obtained with ReporTree v1.0.0 (Mixão et al. 2022) using the 8,558-loci wgMLST profiles of the 1,434 isolates as input and setting distinct “--site-inclusion” thresholds: 0.95, 0.98 and 1.0 (i.e., keep schema loci called in at least 95%, 98% and 100% of the samples, resulting in a 3,261-loci, 3,179-loci and 874-loci allelic matrices, respectively). <strong>Acknowledgements</strong> We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.
## 数据集 本数据集包含1434株**肠炎沙门氏菌(*Salmonella enterica*)**样本的基因组组装结果及其对应的8558位点全基因组(wg)多位点序列分型(Multiple Locus Sequence Type, MLST)图谱。该分型图谱采用chewie-NS数据库(Mamede et al. 2022)中收录的INNUENDO分型方案(Llarena et al. 2018)生成。所有样本均从分析起始阶段(2021年11月)公开于欧洲核苷酸档案库(European Nucleotide Archive, ENA)或美国国家生物技术信息中心(National Center for Biotechnology Information, NCBI)序列读取档案(Sequence Read Archive, SRA)的全基因组测序(Whole-Genome Sequencing, WGS)数据中筛选获得。 本数据集的样本经过严格筛选,以覆盖广泛的遗传多样性(以血清型作为评估维度)。数据集中共涵盖125种不同的沙门氏菌血清型,其中鼠伤寒沙门氏菌(包括单相型)、肠炎沙门氏菌和婴儿沙门氏菌为优势血清型,三者合计占数据集总量的56.2%。 文件"Se_metadata.xlsx"包含每株分离株的元数据信息,具体包括ENA/SRA登录号、生物项目(BioProject)编号、计算机预测MLST型别(ST)以及血清型。"assemblies/"目录下存储了元数据文件中所有分离株的基因组组装结果(格式为.fasta)。文件"profiles/Se_profiles_wgMLST.tsv"为制表符分隔文件,包含元数据文件中所有分离株的8558位点wgMLST分型图谱。文件"profiles/Se_profiles_cgMLST_95.tsv"、"profiles/Se_profiles_cgMLST_98.tsv"和"profiles/Se_profiles_cgMLST_100.tsv"分别对应元数据文件中所有分离株的3261位点、3179位点和874位点核心基因组MLST(cgMLST)分型图谱,上述分型图谱的生成方式详见下文。 ## 数据集筛选与整理 为构建具有遗传多样性的肠炎沙门氏菌基因组组装数据集,本研究收集了本分析起始阶段(2021年11月)Enterobase数据库中公开的分离株遗传多样性(血清型)信息,并整合了既往相关研究的数据。基于上述信息,我们初步筛选得到包含1779个样本的初始数据集,这些样本关联4个生物项目(PRJEB16326、PRJEB20997、PRJEB30335和PRJEB39988)。我们使用fastq-dl v1.0.6从ENA/SRA下载这些样本的WGS数据,并通过Aquamis v1.3.9(Deneke et al. 2021)工具以默认参数完成测序读段的质量控制、修剪及基因组组装。随后使用同一流程完成组装质量控制(QC,包括污染评估)以及MLST ST型别确定。所有通过质量控制的基因组组装结果均被纳入最终数据集。 我们注意到,相当比例的组装结果仅因"NumContamSNVs"参数被标记为"QC失败",提示该参数的阈值设置可能过于严格。在随机抽取子集进行人工检查后,我们将其中对应正确物种的读段占比>98%的组装结果恢复并纳入最终数据集(元数据文件中已对这些样本进行标注)。最终共有1434株分离株通过该整理步骤,被纳入本数据集。我们使用SeqSero2 v1.2.1(Zhang et al. 2019)完成计算机预测血清分型。 本数据集所有分离株的wgMLST分型图谱通过chewBBACA v2.8.5(Silva et al. 2018)工具生成,采用2022年5月31日从chewie-NS数据库下载的8558位点INNUENDO分型方案(Llarena et al. 2018; Mamede et al. 2022)。我们以1434株分离株的8558位点wgMLST分型图谱作为输入,通过ReporTree v1.0.0(Mixão et al. 2022)工具设置不同的"--site-inclusion"阈值(0.95、0.98和1.0,即保留在至少95%、98%和100%的样本中被成功分型的位点),分别得到3261位点、3179位点和874位点的cgMLST分型图谱。 ## 致谢 本研究感谢葡萄牙国家分布式计算基础设施(INCD)提供基因组组装所需的计算资源。INCD由葡萄牙科学技术基金会(FCT)和欧洲区域发展基金(FEDER)资助,项目编号为22153-01/SAICT/2016。



