The OHEJP BeONE Project – Salmonella enterica genome assembly dataset
收藏资源简介:
<strong>Dataset</strong> This dataset comprises the genome assemblies of 1,540 <em>Salmonella enterica</em> samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (https://onehealthejp.eu/jrp-beone/). Additionally, a complementary dataset is also made available (https://zenodo.org/record/7119735), comprising genome assemblies of 1,434 <em>S. enterica</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA). File “<strong>BeONE_Se_metadata.xls</strong>x” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers, in-silico Multi Locus Sequence Type and Serotype, and information regarding year of sampling, country and source. The archive “<strong>BeONE_Se_assemblies.zi</strong>p” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. <strong>Dataset selection and curation</strong> This anonymized dataset of <em>S. enterica</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number PRJEB57179. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (Deneke et al. 2021) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 1,540 isolates passed the dataset curation step and were included in the final dataset. In-silico serotyping was performed with SeqSero2 v1.2.1 (Zhang et al. 2019). <strong>Funding</strong> This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. <strong>Acknowledgements</strong> We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.
**数据集** 本数据集包含BeONE联盟代表全健康欧洲联合计划"BeONE:构建一体化全健康监测工具"(https://onehealthejp.eu/jrp-beone/)所收集的1540株肠炎沙门氏菌(*Salmonella enterica*)样本的基因组组装序列。此外,另有一套补充数据集公开于(https://zenodo.org/record/7119735),包含从欧洲核苷酸档案库(European Nucleotide Archive, ENA)或美国国家生物技术信息中心(National Center for Biotechnology Information, NCBI)序列读取档案(Sequence Read Archive, SRA)的公开全基因组测序(Whole-Genome Sequencing, WGS)数据中筛选出的1434株肠炎沙门氏菌(*S. enterica*)样本的基因组组装序列。文件"**BeONE_Se_metadata.xlsx**"包含各分离株的基因组组装统计信息,涵盖ENA登录号、电子多位点序列分型、血清型,以及采样年份、来源国与分离来源等信息。压缩包"**BeONE_Se_assemblies.zi**"包含元数据文件中所列全部分离株的基因组组装文件(格式为.fasta)。 **数据集筛选与整理** 本匿名肠炎沙门氏菌基因组组装数据集,基于BeONE联盟收集的下一代测序(Next Generation Sequencing, NGS)数据生成,相关数据存于ENA的BioProject登录号PRJEB57179下。读长质量控制、修剪及组装流程均采用Aquamis v1.3.9(Deneke等,2021),并使用默认参数。组装质量控制(QC)包含污染评估与多位点序列分型(MLST)分型,均通过同一流程完成。所有通过QC的基因组组装均被纳入最终数据集。研究过程中发现,相当比例的组装仅因"NumContamSNVs"参数被标记为"QC失败",提示该参数阈值设置可能过于严苛。经随机子集手动检视后,将对应正确物种的读长占比>98%的组装恢复并纳入最终数据集(此类样本已在元数据文件中标注)。最终共有1540株分离株通过数据集整理步骤,被纳入最终数据集。电子血清分型采用SeqSero2 v1.2.1(Zhang等,2019)完成。 **资助说明** 本研究获欧盟地平线2020研究与创新计划资助,资助协议编号为773830:全健康欧洲联合计划。 **致谢** 感谢葡萄牙国家分布式计算基础设施(INCD)提供运行基因组组装所需的计算资源。INCD由葡萄牙科学技术基金会(FCT)与欧洲区域发展基金(FEDER)通过项目22153-01/SAICT/2016资助。



