The OHEJP BeONE Project – Campylobacter jejuni genome assembly dataset
收藏资源简介:
<strong>Dataset</strong> This dataset comprises the genome assemblies of 610 <em>Campylobacter jejuni </em>samples collected by the BeONE Consortium on behalf of the One Health European Joint Programme “BeONE: Building Integrative Tools for One Health Surveillance” (https://onehealthejp.eu/jrp-beone/). Additionally, a complementary dataset is also made available (https://zenodo.org/record/7120166), comprising genome assemblies of 3,076 <em>C. jejuni</em> samples selected among the Whole-Genome Sequencing (WGS) data publicly available in the European Nucleotide Archive (ENA) or in the National Center for Biotechnology Information (NCBI) Sequence Read Archive (SRA). File “<strong>BeONE_Cj_metadata.xlsx</strong>” contains the genome assembly statistics for each isolate, including European Nucleotide Archive accession numbers and <em>in-silico</em> Multi Locus Sequence Type, and information regarding year of sampling, country and source. The archive “<strong>BeONE_Cj_assemblies.zip</strong>” contains all the genome assemblies (.fasta format) of each isolate presented in the metadata file. <strong>Dataset selection and curation</strong> This anonymized dataset of <em>C. jejuni</em> genome assemblies was generated using Next Generation Sequencing data collected within the BeONE Consortium available at the European Nucleotide Archive under BioProject Accession Number PRJEB57119. Read quality control, trimming and assembly were performed with Aquamis v1.3.9 (Deneke et al. 2021) using default parameters. Assembly quality control (QC), including contamination assessment, as well as MLST ST determination were performed with the same pipeline. All genome assemblies passing the QC were included in the final dataset. Among the others, we noticed that a considerable proportion of assemblies was flagged as “QC fail” exclusively due to the “NumContamSNVs” parameter, suggesting that this setting might have been too strict. After manual inspection of a random subset, assemblies for which the percentage of reads corresponding to the correct species was >98% were recovered and integrated in the final dataset (those samples are labeled in the Metadata file). In total, 610 isolates passed the dataset curation step and were included in the final dataset. <strong>Funding</strong> This work was supported by funding from the European Union’s Horizon 2020 Research and Innovation programme under grant agreement No 773830: One Health European Joint Programme. <strong>Acknowledgements</strong> We thank the National Distributed Computing Infrastructure of Portugal (INCD) for providing the necessary resources to run the genome assemblies. INCD was funded by FCT and FEDER under the project 22153-01/SAICT/2016.
数据集 本数据集包含BeONE联盟代表One Health欧洲联合项目“BeONE: 构建综合工具以支撑One Health监测”(https://onehealthejp.eu/jrp-beone/)所收集的610株空肠弯曲菌(Campylobacter jejuni)的基因组组装序列。此外,同步公开了补充数据集(https://zenodo.org/record/7120166),该数据集包含从欧洲核苷酸档案馆(European Nucleotide Archive,ENA)或美国国家生物技术信息中心(National Center for Biotechnology Information,NCBI)序列读取档案(Sequence Read Archive,SRA)公开的全基因组测序(Whole-Genome Sequencing,WGS)数据中筛选出的3076株空肠弯曲菌(C. jejuni)的基因组组装序列。 文件"BeONE_Cj_metadata.xlsx"收录了每份分离株的基因组组装统计信息,包括欧洲核苷酸档案馆收录号、计算机模拟(in-silico)多位点序列分型(Multi Locus Sequence Type,MLST)结果,以及采样年份、分离国家与来源等信息。压缩包"BeONE_Cj_assemblies.zip"包含元数据文件中列出的所有分离株的基因组组装序列(格式为.fasta)。 数据集筛选与整理 本匿名化空肠弯曲菌基因组组装数据集基于BeONE联盟收集的下一代测序(Next Generation Sequencing,NGS)数据生成,相关数据已提交至欧洲核苷酸档案馆,BioProject收录号为PRJEB57119。测序读段的质量控制、修剪及基因组组装均通过Aquamis v1.3.9工具(Deneke等,2021)以默认参数完成。组装质量控制(QC)及污染评估、多位点序列分型(MLST)结果判定均通过同一流程实现。所有通过质量控制的基因组组装序列均纳入最终数据集。 研究过程中发现,相当比例的组装序列仅因“NumContamSNVs”参数被判定为“QC不合格”,提示该参数设置可能过于严苛。经随机抽取样本人工复核后,匹配目标物种的读段占比>98%的组装序列被重新纳入最终数据集(此类样本在元数据文件中已标注)。最终共有610株分离株通过数据集筛选整理步骤,纳入本数据集。 资助 本研究获得欧盟地平线2020研究与创新计划资助,资助协议编号为773830:One Health欧洲联合项目。 致谢 感谢葡萄牙国家分布式计算基础设施(INCD)提供基因组组装所需的计算资源。INCD由葡萄牙科学技术基金会(FCT)与欧洲区域发展基金(FEDER)共同资助,项目编号为22153-01/SAICT/2016。



