遇见数据集

Evaluating the Impact of Different Sequence Databases on Metaproteome Analysis: Insights from a Lab-Assembled Microbial Mixture

收藏
Figshare2016-01-18 更新2026-04-29 收录
官方服务:

资源简介:

Metaproteomics enables the investigation of the protein repertoire expressed by complex microbial communities. However, to unleash its full potential, refinements in bioinformatic approaches for data analysis are still needed. In this context, sequence databases selection represents a major challenge.This work assessed the impact of different databases in metaproteomic investigations by using a mock microbial mixture including nine diverse bacterial and eukaryotic species, which was subjected to shotgun metaproteomic analysis. Then, both the microbial mixture and the single microorganisms were subjected to next generation sequencing to obtain experimental metagenomic- and genomic-derived databases, which were used along with public databases (namely, NCBI, UniProtKB/SwissProt and UniProtKB/TrEMBL, parsed at different taxonomic levels) to analyze the metaproteomic dataset. First, a quantitative comparison in terms of number and overlap of peptide identifications was carried out among all databases. As a result, only 35% of peptides were common to all database classes; moreover, genus/species-specific databases provided up to 17% more identifications compared to databases with generic taxonomy, while the metagenomic database enabled a slight increment in respect to public databases. Then, database behavior in terms of false discovery rate and peptide degeneracy was critically evaluated. Public databases with generic taxonomy exhibited a markedly different trend compared to the counterparts. Finally, the reliability of taxonomic attribution according to the lowest common ancestor approach (using MEGAN and Unipept software) was assessed. The level of misassignments varied among the different databases, and specific thresholds based on the number of taxon-specific peptides were established to minimize false positives. This study confirms that database selection has a significant impact in metaproteomics, and provides critical indications for improving depth and reliability of metaproteomic results. Specifically, the use of iterative searches and of suitable filters for taxonomic assignments is proposed with the aim of increasing coverage and trustworthiness of metaproteomic data.

宏蛋白质组学(Metaproteomics)可用于探究复杂微生物群落所表达的蛋白质组。然而,为充分发挥其应用潜力,当前仍需优化数据分析所用的生物信息学方法。在此背景下,序列数据库的选择是一项核心挑战。本研究以包含9种不同细菌和真核生物的模拟微生物混合物为研究对象,对其实施鸟枪法宏蛋白质组学分析,以此评估不同数据库在宏蛋白质组学研究中的影响。随后,分别对该模拟微生物混合物及其单株微生物开展下一代测序,以此获取实验源性宏基因组与基因组数据库;将上述数据库与公共数据库(即NCBI、UniProtKB/SwissProt及UniProtKB/TrEMBL,按不同分类学层级进行解析)结合,用于宏蛋白质组数据集的分析。首先,针对全部数据库的肽段鉴定数量与肽段重叠情况开展定量对比分析。结果显示,仅35%的肽段可被所有类型的数据库共同鉴定;此外,相较于通用分类学数据库,属/种特异性数据库可多鉴定出最高达17%的肽段,而宏基因组数据库相较于公共数据库仅实现小幅提升。随后,本研究针对数据库在假发现率与肽段简并性层面的表现展开批判性评估。具备通用分类学属性的公共数据库,其表现趋势与其余数据库存在显著差异。最后,本研究评估了基于最低共同祖先法(使用MEGAN与Unipept软件)开展分类学归属注释的可靠性。不同数据库的分类错误分配程度存在差异,本研究基于分类群特异性肽段数量设定了专属阈值,以最大程度降低假阳性结果的出现。本研究证实,数据库选择对宏蛋白质组学研究具有显著影响,并为提升宏蛋白质组学研究结果的深度与可靠性提供了关键指导依据。具体而言,本研究建议采用迭代搜索策略与适配的分类学归属注释过滤条件,以提升宏蛋白质组学数据的覆盖度与可信度。

创建时间:
2016-01-18
二维码
社区交流群
二维码
科研交流群
商业服务