遇见数据集

Breast cancer protein–protein interaction network derived from differentially expressed genes and topological network analysis

收藏
Zenodo2026-04-29 更新2026-05-26 收录
官方服务:

资源简介:

Breast cancer protein–protein interaction network derived from differentially expressed genes and topological network analysis 2. Description This dataset contains a curated list of differentially expressed genes associated with breast cancer, compiled from multiple published studies. The genes were integrated to construct a protein–protein interaction (PPI) network, which was subsequently expanded using interaction data from the STRING database. The objective of the dataset is to support the analysis of molecular interaction networks involved in breast cancer, allowing the identification of highly connected proteins, hub genes, and key regulatory nodes potentially associated with tumor progression and metastasis. The workflow involved three main stages: Compilation of differentially expressed genes (DEGs) reported in published transcriptomic studies of breast cancer. Expansion of the interaction network using experimentally validated and predicted protein interactions obtained from STRING. Topological network analysis to evaluate structural properties of the breast cancer interaction network. The resulting dataset includes the interaction network structure and node attributes used to calculate topological parameters such as node connectivity, centrality measures, and network organization metrics. This dataset can be used for: network biology studies identification of potential molecular biomarkers systems biology analysis of breast cancer validation of computational network analysis methods 3. Methods The dataset was generated using a systems biology approach integrating transcriptomic data and protein interaction networks. 1. Identification of differentially expressed genes Genes associated with breast cancer were collected from previously published transcriptomic studies reporting differential gene expression in breast tumor samples. Gene identifiers were standardized using Gene Symbol, Gene ID, and Ensembl identifiers. 2. Network construction The curated gene list was used as input to construct a protein–protein interaction network using interaction data retrieved from the STRING database. Both experimentally validated and predicted protein interactions were considered. The resulting interaction network represents molecular relationships between proteins encoded by the selected genes. 3. Network expansion The network was expanded by incorporating additional interacting proteins identified through STRING in order to obtain a more comprehensive representation of the molecular interaction landscape associated with breast cancer. 4. Topological analysis Topological properties of the network were analyzed to identify key structural features of the system. Metrics evaluated include: node degree network connectivity centrality measures identification of hub nodes interaction density These parameters allow the identification of proteins that may play important regulatory roles in the network. 4. Dataset structure The dataset is provided as an Excel file (.xlsx) containing information about genes and their interactions within the protein–protein interaction network. The dataset includes the following types of information: Gene information Gene symbol Gene ID Ensembl gene ID Source publication where the gene was reported as differentially expressed Interaction network data interacting protein pairs interaction identifiers source of interaction data Network analysis data node identifiers interaction relationships between proteins variables used for calculating topological parameters The dataset represents the interaction edges and node attributes required to reconstruct the protein interaction network used in the study.

提供机构:
Zenodo
创建时间:
2026-04-29
二维码
社区交流群
二维码
科研交流群
商业服务