遇见数据集

Fast Computing for Distance Covariance

收藏
NIAID Data Ecosystem2026-03-08 收录
官方服务:

资源简介:

Distance covariance and distance correlation have been widely adopted in measuring dependence of a pair of random variables or random vectors. If the computation of distance covariance and distance correlation is implemented directly accordingly to its definition then its computational complexity is O(n2) which is a disadvantage compared to other faster methods. In this paper we show that the computation of distance covariance and distance correlation of real valued random variables can be implemented by an O(nlog n) algorithm and this is comparable to other computationally efficient algorithms. The new formula we derive for an unbiased estimator for squared distance covariance turns out to be a U-statistic. This fact implies some nice asymptotic properties that were derived before via more complex methods. We apply the fast computing algorithm to some synthetic data. Our work will make distance correlation applicable to a much wider class of problems. A supplementary file to this article includes a Matlab and C based software that realizes the proposed algorithm.

距离协方差(distance covariance)与距离相关系数(distance correlation)已被广泛应用于衡量一对随机变量或随机向量之间的相依性。若直接依据定义实现距离协方差与距离相关系数的计算,其计算复杂度为O(n²),相较于其他更高效的计算方法,该复杂度是其明显短板。本文证明,针对实值随机变量的距离协方差与距离相关系数的计算,可通过O(n log n)复杂度的算法实现,该算法的计算效率可与其他同类高效计算算法相媲美。我们推导得到的平方距离协方差无偏估计量的新公式,本质上属于U统计量(U-statistic)。这一性质揭示了若干优良的渐近特性,而这些特性此前需通过更为复杂的方法方可推导得出。我们将该快速计算算法应用于若干合成数据集。本研究将使距离相关系数可适用于更广泛的问题范畴。本文的补充文件包含了基于Matlab与C语言实现所提算法的软件程序。

创建时间:
2015-06-18
二维码
社区交流群
二维码
科研交流群
商业服务