遇见数据集

Whole proteome-level GPS predictions (part 2)

收藏
Mendeley Data2024-01-31 更新2024-06-28 收录
官方服务:

资源简介:

This is the expanded set of all predictions for GPS, run on the entire reference proteome, including sites not known to be phosphorylated. This dataset is used to perform a fast update when new phosphosites are discovered. The uncompressed folder will yield a large CSV file with predictions in list format (i.e. one line per kinase-substrate prediction) Columns in this order:substrate_id - unique substrate (accession_site) IDsubstrate_acc - Uniprot accession of substrate proteinsubstrate_name - Name of proteinsite - amino acid type and position (S5, means serine position 5)pep - 15-amino acid sequence centered on site of phosphorylationscore - prediction algorithm scoreKinase Name - name of kinase by our controlled ontology (found in this project) Each entry indicates the protein at position (identified by peptide) and has a score weight prediction for the given kinase. FOR FULL GPS RAW, you must combine this with the first zip part. Please be sure to download both into the same directory before unzipping. The final, uncompressed file, is 24GB.

本数据集为针对GPS的全量预测结果扩展集,基于完整参考蛋白质组运行生成,涵盖尚未被证实为磷酸化位点的序列位点。该数据集可用于在发现新磷酸化位点时执行快速更新。未压缩的文件夹将生成一个大型CSV文件,其中的预测结果以列表格式存储(即每条激酶-底物预测对应一行)。各列按如下顺序排列:substrate_id(底物唯一(登录号_位点)标识符)、substrate_acc(底物蛋白的UniProt登录号)、substrate_name(蛋白名称)、site(氨基酸类型与位点位置,例如S5代表第5位丝氨酸)、pep(以磷酸化位点为中心的15个氨基酸序列)、score(预测算法得分)、Kinase Name(本项目受控本体论所定义的激酶名称,可在本项目中查阅)。每条条目均代表由肽段确定的位置处的蛋白质,并针对指定激酶给出了得分权重预测。若需获取完整的GPS原始数据,需将本数据集与第一个压缩分卷进行合并。请确保将两个压缩包下载至同一目录后再执行解压缩操作。最终解压缩后的文件大小为24GB。

创建时间:
2024-01-31
搜集汇总
数据集介绍
Whole proteome-level GPS predictions (part 2) 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务