XMAn_Homo_Sapiens_Mutated_Peptides_DB_Disease_fasta
收藏资源简介:
To enable the identification of mutated peptide sequences in complex biological samples, in this work, a disease protein database with mutation information collected from several public resources such as OMIM and UniProtKB, was developed. In-house developed Perl-scripts were used to search and process the data, and to translate each gene-level mutation into a mutated peptide sequence. The disease mutation database comprises a total of 27,148 peptide entries from 2913 protein IDs. A description line for each entry provides the parent protein ID and name, the cDNA- and protein-level mutation site and type, the originating database, and the tissue type and corresponding hits. The database is FASTA formatted to enable data retrieval by commonly used tandem MS search engines.
为实现复杂生物样本中突变肽序列的鉴定,本研究构建了一款收录OMIM(在线人类孟德尔遗传数据库,Online Mendelian Inheritance in Man)、UniProtKB(通用蛋白质知识库,Universal Protein Knowledgebase)等多个公共数据源突变信息的疾病蛋白质数据库。本研究自主开发的Perl脚本用于数据的检索与处理,并将每个基因层面的突变转换为突变肽序列。该疾病突变数据库共计包含2913个蛋白质ID对应的27148条肽条目。每条条目均配有描述行,涵盖其父蛋白质ID与名称、cDNA(互补DNA,complementary DNA)及蛋白质层面的突变位点与突变类型、来源数据库,以及组织类型与对应匹配命中数。该数据库采用FASTA格式构建,可通过常用的串联质谱(tandem MS,tandem mass spectrometry)搜索引擎实现数据检索。



