Population structure and gene flux of Listeria monocytogenes ST121 reveal prophages as a candidate driver of adaptation and persistence in food production environments
收藏资源简介:
Background Listeria monocytogenes is a bacterial pathogen found in an increasing number of food categories, potentially reflecting an expanding niche and food safety risk profile. In the United Kingdom, L. monocytogenes sequence type (ST) 121 is more frequently isolated from foods and food environments than from cases of clinical listeriosis, consistent with a relatively low pathogenicity. In this study, we determined the microevolution associated with the environmental persistence of a L. monocytogenes clone by investigating clone-specific genome features in the context of the ST121 population structure from international sources. To enable unambiguous comparative genomic analysis of ST121 strains we constructed 16 new high-quality genome assemblies from L. monocytogenes isolated from foods, food environments, and human-clinical sources in the UK from 1987 to 2019. Genome assemblies Hybrid assemblies of strains 07-032, BL87-028, BL89-019, BL89-020, 399454, 246242, 163381, 241745, 457441, 389690, 535305, 396044, 246255, 259390, 754322 and 757923 were produced using long and short reads. Long reads were adapter trimmed with Porechop v0.2.3 and filtered with filtlong v0.2.1 to keep the 95% of the best reads longer than 6000 bp. Contamination of the filtered reads was assessed using kraken2 v2.1.1 against the kraken2 database Standard. A long-read assembly was produced using the filtlong filtered reads as input to the Flye assembler v2.7 with the nano-raw mode. Long reads were filtered again with filtlong to keep reads longer than 1000 bp, and the resulting reads were used to polish the circular long-read assemblies using medaka v1.4.3. Paired-end short reads were filtered with trimmomatic v0.38.0 using default parameters with a sliding window of 4 bp and quality of 20 and inspected for contamination using kraken2 v2.1.1. The long-read assembly was further polished with the filtered short reads using pilon v1.24 with the following nextflow script. Using default parameters, Fastq files from strains 1537877 and 1594542 underwent trimming using TrimGalore v0.4.3. Next, the quality of the resulting paired-end sequences was assessed using FastQC v0.11.8, sequences were filtered based on a GC content above 39% and N bases in a proportion higher than 5%. Assemblies were constructed using shovill v1.1.0 with SPAdes v3.14 as assembler. Assemblies of strains H094800054 and 130 were downloaded from the Institute Pasteur's BIGSdb on 22 September 2022. Genomes were annotated using Bakta v1.6.1. Files File 07_032.fsa corresponds to the hybrid assembly of L. monocytogenes strain 07-032 and annotations are in file 07_032.gff3File BL87_028.fsa corresponds to the hybrid assembly of L. monocytogenes strain BL87-028 and annotations are in file BL87_028.gff3File BL89_019.fsa corresponds to the hybrid assembly of L. monocytogenes strain BL89-019 and annotations are in file BL89_019.gff3File BL89_020.fsa corresponds to the hybrid assembly of L. monocytogenes strain BL89-020 and annotations are in file BL89_020.gff3File 163381_SRR6807418.fsa corresponds to the hybrid assembly of L. monocytogenes strain 163381 and annotations are in file 163381_SRR6807418.gff3File 399454_SRR7163869.fsa corresponds to the hybrid assembly of L. monocytogenes strain 399454 and annotations are in file 399454_SRR7163869.gff3File 241745_SRR7167591.fsa corresponds to the hybrid assembly of L. monocytogenes strain 241745 and annotations are in file 241745_SRR7167591.gff3File 457441_SRR7827106.fsa corresponds to the hybrid assembly of L. monocytogenes strain 457441 and annotations are in file 457441_SRR7827106.gff3File 389690_SRR7841399.fsa corresponds to the hybrid assembly of L. monocytogenes strain 389690 and annotations are in file 389690_SRR7841399.gff3File 246242_SRR7850130.fsa corresponds to the hybrid assembly of L. monocytogenes strain 246242 and annotations are in file 246242_SRR7850130.gff3File 535305_SRR7866357.fsa corresponds to the hybrid assembly of L. monocytogenes strain 535305 and annotations are in file 535305_SRR7866357.gff3File 396044_SRR7866629.fsa corresponds to the hybrid assembly of L. monocytogenes strain 396044 and annotations are in file 396044_SRR7866629.gff3File 246255_SRR7873592.fsa corresponds to the hybrid assembly of L. monocytogenes strain 246255 and annotations are in file 246255_SRR7873592.gff3File 259390_SRR7873684.fsa corresponds to the hybrid assembly of L. monocytogenes strain 259390 and annotations are in file 259390_SRR7873684.gff3File 754322_SRR9226492.fsa corresponds to the hybrid assembly of L. monocytogenes strain 754322 and annotations are in file 754322_SRR9226492.gff3File 757923_SRR9298670.fsa corresponds to the hybrid assembly of L. monocytogenes strain 757923 and annotations are in file 757923_SRR9298670.gff3 File 1838_H094800054.gff3 corresponds to the annotation of the draft assembly of L. monocytogenes strain H094800054File 79394_130.gff3 corresponds to the annotation of the draft assembly of L. monocytogenes strain 130File 1537877_SRR17120467.gff3 corresponds to the annotation of the draft assembly of L. monocytogenes strain 1537877File 1594542_SRR18333768.gff3 corresponds to the annotation of the draft assembly of L. monocytogenes strain 1594542 File All_pangenome_reference.fa corresponds to the pangenome sequences of 482 L. monocytogenes strains of ST121File Chr_pangenome_reference.fa corresponds to the chromosome reference-based gene database build from 108 completely assembled L. monocytogenes reference genomesFile Phage_pangenome_reference.fa corresponds to the to the phage reference-based gene database build from 30 Listeria phage sequences obtained from NCBIFile Plasmids_pangenome_reference.fa corresponds to the plasmid reference-based gene database build from 41 Listeria plasmid sequences obtained from NCBI
【背景】 单核细胞增生李斯特菌(Listeria monocytogenes)是一种可在日益增多的食品品类中检出的细菌性病原体,这一现象或反映其生态位扩张与食品安全风险谱的扩展。在英国,相较于临床李斯特菌病病例,序列型(sequence type, ST)121型单核细胞增生李斯特菌更常从食品及食品环境中分离得到,这与其相对较低的致病性相符。本研究通过结合国际来源的ST121种群结构背景,分析克隆特异性基因组特征,解析与单核细胞增生李斯特菌克隆株环境持续定植相关的微进化过程。为实现ST121菌株的无歧义比较基因组分析,我们从1987年至2019年英国分离自食品、食品环境及人类临床样本的单核细胞增生李斯特菌中,构建了16套全新的高质量基因组组装体。 【基因组组装】 对菌株07-032、BL87-028、BL89-019、BL89-020、399454、246242、163381、241745、457441、389690、535305、396044、246255、259390、754322及757923采用长读长与短读长测序数据进行混合组装。首先使用Porechop v0.2.3对长读长序列进行接头修剪,并通过filtlong v0.2.1过滤,保留95%长度大于6000 bp的优质读段。使用Kraken2 v2.1.1结合Kraken2标准数据库,对过滤后的读段进行污染评估。以filtlong过滤后的读段为输入,通过Flye组装器v2.7的纳米原始(nano-raw)模式构建长读长组装体。再次使用filtlong对长读长序列进行过滤,保留长度大于1000 bp的读段,随后利用medaka v1.4.3对环状长读长组装体进行序列纠错抛光。采用Trimmomatic v0.38.0以默认参数对双端短读长序列进行过滤,设置4 bp滑动窗口与质量阈值20,并通过Kraken2 v2.1.1检测污染情况。最后使用Pilon v1.24结合过滤后的短读长序列,通过Nextflow脚本对长读长组装体进行进一步纠错抛光。 针对菌株1537877与1594542的Fastq文件,采用TrimGalore v0.4.3以默认参数进行序列修剪。随后使用FastQC v0.11.8评估所得双端序列的质量,并基于GC含量高于39%、N碱基占比低于5%的标准对序列进行过滤。最终采用Shovill v1.1.0,以SPAdes v3.14作为组装工具构建基因组组装体。 菌株H094800054与130的基因组组装体于2022年9月22日从巴斯德研究所BIGSdb数据库下载得到。 所有基因组均通过Bakta v1.6.1完成注释。 【文件说明】 文件07_032.fsa对应单核细胞增生李斯特菌菌株07-032的混合组装序列,其注释信息存储于文件07_032.gff3中。 文件BL87_028.fsa对应单核细胞增生李斯特菌菌株BL87-028的混合组装序列,其注释信息存储于文件BL87_028.gff3中。 文件BL89_019.fsa对应单核细胞增生李斯特菌菌株BL89-019的混合组装序列,其注释信息存储于文件BL89_019.gff3中。 文件BL89_020.fsa对应单核细胞增生李斯特菌菌株BL89-020的混合组装序列,其注释信息存储于文件BL89_020.gff3中。 文件163381_SRR6807418.fsa对应单核细胞增生李斯特菌菌株163381的混合组装序列,其注释信息存储于文件163381_SRR6807418.gff3中。 文件399454_SRR7163869.fsa对应单核细胞增生李斯特菌菌株399454的混合组装序列,其注释信息存储于文件399454_SRR7163869.gff3中。 文件241745_SRR7167591.fsa对应单核细胞增生李斯特菌菌株241745的混合组装序列,其注释信息存储于文件241745_SRR7167591.gff3中。 文件457441_SRR7827106.fsa对应单核细胞增生李斯特菌菌株457441的混合组装序列,其注释信息存储于文件457441_SRR7827106.gff3中。 文件389690_SRR7841399.fsa对应单核细胞增生李斯特菌菌株389690的混合组装序列,其注释信息存储于文件389690_SRR7841399.gff3中。 文件246242_SRR7850130.fsa对应单核细胞增生李斯特菌菌株246242的混合组装序列,其注释信息存储于文件246242_SRR7850130.gff3中。 文件535305_SRR7866357.fsa对应单核细胞增生李斯特菌菌株535305的混合组装序列,其注释信息存储于文件535305_SRR7866357.gff3中。 文件396044_SRR7866629.fsa对应单核细胞增生李斯特菌菌株396044的混合组装序列,其注释信息存储于文件396044_SRR7866629.gff3中。 文件246255_SRR7873592.fsa对应单核细胞增生李斯特菌菌株246255的混合组装序列,其注释信息存储于文件246255_SRR7873592.gff3中。 文件259390_SRR7873684.fsa对应单核细胞增生李斯特菌菌株259390的混合组装序列,其注释信息存储于文件259390_SRR7873684.gff3中。 文件754322_SRR9226492.fsa对应单核细胞增生李斯特菌菌株754322的混合组装序列,其注释信息存储于文件754322_SRR9226492.gff3中。 文件757923_SRR9298670.fsa对应单核细胞增生李斯特菌菌株757923的混合组装序列,其注释信息存储于文件757923_SRR9298670.gff3中。 文件1838_H094800054.gff3对应单核细胞增生李斯特菌菌株H094800054草图组装序列的注释信息。 文件79394_130.gff3对应单核细胞增生李斯特菌菌株130草图组装序列的注释信息。 文件1537877_SRR17120467.gff3对应单核细胞增生李斯特菌菌株1537877草图组装序列的注释信息。 文件1594542_SRR18333768.gff3对应单核细胞增生李斯特菌菌株1594542草图组装序列的注释信息。 文件All_pangenome_reference.fa对应482株ST121型单核细胞增生李斯特菌的泛基因组序列。 文件Chr_pangenome_reference.fa对应基于108株完整组装的单核细胞增生李斯特菌参考基因组构建的染色体参考基因数据库。 文件Phage_pangenome_reference.fa对应从NCBI获取的30株李斯特菌噬菌体序列构建的噬菌体参考基因数据库。 文件Plasmids_pangenome_reference.fa对应从NCBI获取的41株李斯特菌质粒序列构建的质粒参考基因数据库。



