On Simulating Skewed and Cluster-Weighted Data for Studying Performance of Clustering Algorithms
收藏资源简介:
In this article, extensions to the recently introduced concept of pairwise overlap between mixture components are proposed. The notion of overlap is useful for studying the systematic performance of clustering algorithms. Existing methods can be used for simulating elliptical data according to pre-specified overlap characteristics. First, an approach to simulating skewed clusters with a desired overlap is proposed. Next, an extension to measuring overlap in cluster-weighted models is considered. Thus, this article provides important extensions to the existing methods for simulating heterogeneous data for studying the systematic performance of clustering algorithms. Supplementary materials for this article are available online.
本文针对近年提出的混合分量间成对重叠(pairwise overlap)概念提出了拓展方法。重叠这一概念对于研究聚类算法的系统性表现具有重要价值。现有方法可依据预设的重叠特征,实现椭圆分布数据的模拟生成。首先,本文提出了一种可生成具有目标重叠程度的偏态聚类簇的模拟方法;其次,针对聚类加权模型(cluster-weighted models)中的重叠度度量方法提出了拓展。综上,本文为现有用于模拟异质数据以研究聚类算法系统性表现的方法提供了重要拓展。本文的补充材料可在线获取。



