遇见数据集

Selection of the Regularization Parameter in Graphical Models Using Network Characteristics

收藏
Figshare2017-08-14 更新2026-04-29 收录
官方服务:

资源简介:

Gaussian graphical models represent the underlying graph structure of conditional dependence between random variables, which can be determined using their partial correlation or precision matrix. In a high-dimensional setting, the precision matrix is estimated using penalized likelihood by adding a penalization term, which controls the amount of sparsity in the precision matrix and totally characterizes the complexity and structure of the graph. The most commonly used penalization term is the L1 norm of the precision matrix scaled by the regularization parameter, which determines the trade-off between sparsity of the graph and fit to the data. In this article, we propose several procedures to select the regularization parameter in the estimation of graphical models that focus on recovering reliably the appropriate network structure of the graph. We conduct an extensive simulation study to show that the proposed methods produce useful results for different network topologies. The approaches are also applied in a high-dimensional case study of gene expression data with the aim to discover the genes relevant to colon cancer. Using these data, we find graph structures, which are verified to display significant biological gene associations. Supplementary material is available online.

高斯图模型(Gaussian graphical models)刻画了随机变量间条件依赖关系的底层图结构,该结构可通过变量的偏相关系数或精度矩阵(precision matrix)加以确定。在高维场景下,精度矩阵的估计需通过添加惩罚项实现惩罚似然估计,该惩罚项可控制精度矩阵的稀疏程度,并全面刻画该图的复杂度与结构。最常用的惩罚项为经正则化参数(regularization parameter)缩放后的精度矩阵L1范数,该参数可权衡图结构的稀疏性与数据拟合度。本文提出了多种用于图模型估计中正则化参数选择的方法,其核心目标为可靠地恢复出图的恰当网络结构。我们开展了大规模模拟研究,结果表明所提方法在多种网络拓扑结构下均可取得良好效果。我们还将所提方法应用于一项高维基因表达数据案例研究中,旨在发掘与结肠癌相关的基因。基于该数据集,我们得到了图结构,经验证该结构可体现出具有显著生物学意义的基因关联。本文补充材料可在线获取。

创建时间:
2017-08-14
二维码
社区交流群
二维码
科研交流群
商业服务