XMAn: A <i>Homo sapiens</i> Mutated-Peptide Database for the MS Analysis of Cancerous Cell States
收藏资源简介:
To enable the identification of mutated peptide sequences in complex biological samples, in this work, two novel cancer- and disease-related protein databases with mutation information collected from several public resources such as COSMIC, IARC P53, OMIM, and UniProtKB were developed. In-house developed Perl scripts were used to search and process the data and to translate each gene-level mutation into a mutated peptide sequence. The cancer and disease mutation databases comprise a total of 872 125 and 27 148 peptide entries from 25 642 and 2913 proteins, respectively. A description line for each entry provides the parent protein ID and name, the cDNA- and protein-level mutation site and type, the originating database, and the disease or cancer tissue type and corresponding hits. The two databases are FASTA-formatted to enable data retrieval by commonly used tandem MS search engines. While the largest number of mutations were encountered for the amino acids A/D/E/G/L/P/R/S, the global mutation profiles replicate closely the outcome of the 1000 Genomes Project aimed at cataloguing natural mutations in the human population. The affected proteins were primarily involved in transcription regulation, splicing, protein synthesis/folding/binding, redox/energy production, adhesion/motility, and to some extent in DNA damage repair and signaling. The applicability of the database to identifying the presence of mutated peptides was investigated with MCF-7 breast cancer cell extracts.
为实现复杂生物样本中突变肽序列的鉴定,本研究构建了两款全新的癌症及疾病相关蛋白质数据库,其突变信息取自COSMIC、IARC P53、OMIM、UniProtKB等多个公开数据源。本研究使用自研Perl脚本完成数据的检索与处理,并将每一处基因层面的突变转换为突变肽序列。该癌症与疾病突变数据库分别包含来自25642个和2913个蛋白质的872125条、27148条肽段条目。每条条目均附带描述信息,涵盖父蛋白ID与名称、互补DNA(cDNA)及蛋白质层面的突变位点与突变类型、来源数据库、疾病或癌症组织类型及对应匹配结果。两款数据库均采用FASTA格式,可通过常用的串联质谱搜索引擎进行数据检索。尽管突变频次最高的氨基酸为丙氨酸(A)、天冬氨酸(D)、谷氨酸(E)、甘氨酸(G)、亮氨酸(L)、脯氨酸(P)、精氨酸(R)、丝氨酸(S),但其整体突变谱与旨在收录人类群体自然突变的千人基因组计划(1000 Genomes Project)结果高度吻合。受突变影响的蛋白质主要参与转录调控、剪接、蛋白质合成/折叠/结合、氧化还原/能量产生、黏附/运动等生物学过程,在一定程度上也涉及脱氧核糖核酸(DNA)损伤修复与信号传导通路。本研究以MCF-7乳腺癌细胞提取物为样本,验证了该数据库在鉴定突变肽序列方面的适用性。



