遇见数据集

Interacting networks of resistance, virulence and core machinery genes identified by genome-wide epistasis analysis

收藏
Figshare2017-02-17 更新2026-04-29 收录
官方服务:

资源简介:

Recent advances in the scale and diversity of population genomic datasets for bacteria now provide the potential for genome-wide patterns of co-evolution to be studied at the resolution of individual bases. Here we describe a new statistical method, genomeDCA, which uses recent advances in computational structural biology to identify the polymorphic loci under the strongest co-evolutionary pressures. We apply genomeDCA to two large population data sets representing the major human pathogens Streptococcus pneumoniae (pneumococcus) and Streptococcus pyogenes (group A Streptococcus). For pneumococcus we identified 5,199 putative epistatic interactions between 1,936 sites. Over three-quarters of the links were between sites within the pbp2x, pbp1a and pbp2b genes, the sequences of which are critical in determining non-susceptibility to beta-lactam antibiotics. A network-based analysis found these genes were also coupled to that encoding dihydrofolate reductase, changes to which underlie trimethoprim resistance. Distinct from these antibiotic resistance genes, a large network component of 384 protein coding sequences encompassed many genes critical in basic cellular functions, while another distinct component included genes associated with virulence. The group A Streptococcus (GAS) data set population represents a clonal population with relatively little genetic variation and a high level of linkage disequilibrium across the genome. Despite this, we were able to pinpoint two RNA pseudouridine synthases, which were each strongly linked to a separate set of loci across the chromosome, representing biologically plausible targets of co-selection. The population genomic analysis method applied here identifies statistically significantly co-evolving locus pairs, potentially arising from fitness selection interdependence reflecting underlying protein-protein interactions, or genes whose product activities contribute to the same phenotype. This discovery approach greatly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for targeted experimental work.

近年来,细菌群体基因组数据集(population genomic datasets)在规模与多样性方面取得的进展,使得我们能够以单碱基分辨率对全基因组共进化模式展开研究。本文介绍一种全新的统计方法genomeDCA,该方法借助计算结构生物学(computational structural biology)领域的最新进展,识别受最强共进化选择压力作用的多态位点(polymorphic loci)。我们将genomeDCA应用于两组大型群体数据集,这两组数据集分别对应两类主要的人类致病菌:肺炎链球菌(*Streptococcus pneumoniae*,俗称肺炎球菌pneumococcus)以及化脓性链球菌(*Streptococcus pyogenes*,即A群链球菌group A Streptococcus)。针对肺炎球菌,我们在1936个位点间识别出5199组推定的上位相互作用(epistatic interactions)。其中超过四分之三的相互作用链接位于pbp2x、pbp1a与pbp2b基因内部的位点之间,这些基因的序列是决定菌株对β-内酰胺类抗生素不敏感性的关键因素。基于网络的分析显示,这些基因同时与编码二氢叶酸还原酶(dihydrofolate reductase)的基因存在关联,该基因的变异是甲氧苄啶耐药性的分子基础。与上述抗生素耐药基因不同,一个包含384个蛋白质编码序列的大型网络模块涵盖了诸多参与基础细胞功能的关键基因;而另一个独立的网络模块则包含了与毒力相关的基因。A群链球菌(group A Streptococcus,简称GAS)的群体数据集对应的是一个克隆群体,其遗传变异相对较少,且全基因组范围内存在高水平的连锁不平衡(linkage disequilibrium)。尽管存在上述特征,我们仍成功定位到两个RNA假尿苷合酶(RNA pseudouridine synthases)基因,二者分别与染色体上的一组独立位点存在强关联,这些关联位点均为符合生物学合理性的共选择靶标。本文所采用的群体基因组分析方法,可识别出具有统计学显著性的共进化位点对,这类位点对可能源于反映底层蛋白质-蛋白质相互作用的适合度选择互作,或是源于编码产物参与同一表型调控的基因间的关联。该发现方法极大地拓展了上位性分析在系统生物学(systems biology)领域的未来应用潜力,同时可作为全基因组关联研究(genome-wide association studies)的补充手段,用于构建靶向实验研究的相关假说。

创建时间:
2017-02-17
二维码
社区交流群
二维码
科研交流群
商业服务