遇见数据集

Data_Sheet_2_Comparative Analysis of Unsupervised Protein Similarity Prediction Based on Graph Embedding.ZIP

收藏
NIAID Data Ecosystem2026-03-12 收录
官方服务:

资源简介:

The study of protein–protein interaction and the determination of protein functions are important parts of proteomics. Computational methods are used to study the similarity between proteins based on Gene Ontology (GO) to explore their functions and possible interactions. GO is a series of standardized terms that describe gene products from molecular functions, biological processes, and cell components. Previous studies on assessing the similarity of GO terms were primarily based on Information Content (IC) between GO terms to measure the similarity of proteins. However, these methods tend to ignore the structural information between GO terms. Therefore, considering the structural information of GO terms, we systematically analyze the performance of the GO graph and GO Annotation (GOA) graph in calculating the similarity of proteins using different graph embedding methods. When applied to the actual Human and Yeast datasets, the feature vectors of GO terms and proteins are learned based on different graph embedding methods. To measure the similarity of the proteins annotated by different GO numbers, we used Dynamic Time Warping (DTW) and cosine to calculate protein similarity in GO graph and GOA graph, respectively. Link prediction experiments were then performed to evaluate the reliability of protein similarity networks constructed by different methods. It is shown that graph embedding methods have obvious advantages over the traditional IC-based methods. We found that random walk graph embedding methods, in particular, showed excellent performance in calculating the similarity of proteins. By comparing link prediction experiment results from GO(DTW) and GOA(cosine) methods, it is shown that GO(DTW) features provide highly effective information for analyzing the similarity among proteins.

蛋白质-蛋白质相互作用研究与蛋白质功能鉴定是蛋白质组学的重要组成部分。研究人员常基于基因本体论(Gene Ontology,GO)开展蛋白质间相似性的计算分析,以探索其功能与潜在相互作用。GO是一系列标准化术语,可从分子功能、生物过程与细胞组分三个维度描述基因产物。既往针对GO术语相似性的评估研究,主要依托GO术语间的信息内容(Information Content,IC)来衡量蛋白质相似性,但这类方法往往忽略了GO术语间的结构信息。因此,本研究考虑GO术语的结构信息,系统分析了不同图嵌入方法在基于GO图与基因本体注释(GO Annotation,GOA)图计算蛋白质相似性时的表现。在应用于实际的人类与酵母数据集时,我们基于不同图嵌入方法学习得到了GO术语与蛋白质的特征向量。为衡量不同GO编号注释的蛋白质间相似性,我们分别在GO图与GOA图中采用动态时间规整(Dynamic Time Warping,DTW)与余弦相似度计算蛋白质相似性。随后通过链接预测实验,评估了不同方法构建的蛋白质相似性网络的可靠性。结果表明,图嵌入方法相较传统基于IC的方法具有显著优势,其中随机游走类图嵌入方法在蛋白质相似性计算中表现尤为出色。通过对比GO(DTW)与GOA(cosine)方法的链接预测实验结果,可发现GO(DTW)特征可为蛋白质相似性分析提供高效的信息支撑。

创建时间:
2021-09-22
二维码
社区交流群
二维码
科研交流群
商业服务