VA-Index: Quantifying Assortativity Patterns in Networks with Multidimensional Nodal Attributes
收藏资源简介:
Network connections have been shown to be correlated with structural or external attributes of the network vertices in a variety of cases. Given the prevalence of this phenomenon network scientists have developed metrics to quantify its extent. In particular, the assortativity coefficient is used to capture the level of correlation between a single-dimensional attribute (categorical or scalar) of the network nodes and the observed connections, i.e., the edges. Nevertheless, in many cases a multi-dimensional, i.e., vector feature of the nodes is of interest. Similar attributes can describe complex behavioral patterns (e.g., mobility) of the network entities. To date little attention has been given to this setting and there has not been a general and formal treatment of this problem. In this study we develop a metric, the vector assortativity index (VA-index for short), based on network randomization and (empirical) statistical hypothesis testing that is able to quantify the assortativity patterns of a network with respect to a vector attribute. Our extensive experimental results on synthetic network data show that the VA-index outperforms a baseline extension of the assortativity coefficient, which has been used in the literature to cope with similar cases. Furthermore, the VA-index can be calibrated (in terms of parameters) fairly easy, while its benefits increase with the (co-)variance of the vector elements, where the baseline systematically over(under)estimate the true mixing patterns of the network.
大量研究表明,在诸多实际场景中,网络连接与网络节点的结构属性或外部属性存在相关性。鉴于该现象的普遍性,网络科学研究者已开发出多种指标以量化其关联程度。具体而言,同配性系数(assortativity coefficient)可用于衡量网络节点的单维属性(分类属性或标量属性)与观测到的连边(即网络边)之间的相关水平。然而,在诸多研究场景中,研究者往往关注节点的多维属性,即向量特征。这类属性可用于描述网络实体的复杂行为模式(如移动性)。迄今为止,针对该类场景的相关研究尚少,尚未有针对该问题的通用且规范的处理框架。本研究基于网络随机化方法与(经验)统计假设检验,提出了向量同配性指数(vector assortativity index,简称VA-index)这一指标,可用于量化网络针对向量属性的同配模式。我们在合成网络数据集上开展的大量实验结果表明,VA-index的性能优于现有文献中用于处理同类场景的同配性系数基线扩展方法。此外,VA-index的参数校准过程较为简便,且其性能优势会随向量元素的(共)方差增大而愈发显著;而该基线方法则会系统性地高估或低估网络的真实混合模式。




