A massive proteogenomic screen identifies thousands of novel human protein coding sequences
收藏资源简介:
Accurate annotation of genes in the human genome is fundamental for biomedical research and genomic data interpretation. The Ensembl, RefSeq, and GENCODE consortiums continuously update the human genome annotations based on new computational and experimental evidence, and new proteins were identified constantly. The Genotype-Tissue Expression (GTEx) project has generated more than 15,000 RNA sequencing dataset from multiple-tissues of more than 800 donors which allows to model almost all transcripts and proteins in the human genome. Using proteins translated from the GTEx transcript model, more than 21 million in-silico trypsin-digested peptides were generated. To identify high-confidence novel proteins with proteomic support, we screened more than 2,000 proteomic projects in the PRIDE database and selected more than 50,000 mass spectrometry (MS) runs from 923 projects. These MS data were used to validate the predicted novel peptides. With a stringent standard, we identified almost 20,000 novel peptides. This dataset include files used in the the above analysis. More details can be found in the GitHub page (https://github.com/ATPs/human_novo_protein_2021).
精准注释人类基因组中的基因,是生物医学研究与基因组数据解析的核心基础。Ensembl、RefSeq与GENCODE三大联盟会基于最新的计算与实验证据持续更新人类基因组注释信息,且新蛋白质的鉴定工作也在持续推进。基因型-组织表达(Genotype-Tissue Expression, GTEx)项目已从800余名供体的多类组织中生成了超过15000套RNA测序数据集,借此可构建人类基因组中几乎所有转录本与蛋白质的模型。基于GTEx转录本模型翻译得到的蛋白质,我们共生成了超过2100万条虚拟胰蛋白酶消化肽段(in-silico trypsin-digested peptides)。为鉴定具备蛋白质组学证据支持的高可信度新型蛋白质,我们对PRIDE数据库中的2000余项蛋白质组学项目进行了筛选,最终从923个项目中选取了超过50000次质谱(mass spectrometry, MS)检测数据。上述质谱数据被用于验证预测得到的新型肽段。通过严格的筛选标准,我们最终鉴定出了近20000条新型肽段。本数据集包含上述分析过程中使用的全部文件。更多详细信息可参阅其GitHub页面:https://github.com/ATPs/human_novo_protein_2021。



