遇见数据集

Deep Clustering Evaluation: How to Validate Internal Clustering Validation Measures

收藏
Figshare2025-12-05 更新2026-04-28 收录
官方服务:

资源简介:

Deep clustering partitions complex high-dimensional data using deep neural networks for clustering. It involves projecting data into lower-dimensional embeddings before partitioning, which embarks unique evaluation challenges. Traditional clustering validation measures, designed for low-dimensional spaces, are problematic for deep clustering for two reasons: (a) the curse of dimensionality when applied to the high-dimensional input data, and (b) unreliable comparison of clustering results when applied to embedded data from different embedding spaces, owing to variations in training procedures and model parameter settings. This article addresses these unresolved and often overlooked challenges in evaluating clustering within deep learning. We propose a systematic evaluation framework for internal clustering validation measures that: (a) theoretically establishes why traditional measures are ineffective when applied to input data or across disparate embedding spaces paired with partitioning outcomes; (b) identifies embedding spaces that endorse reliable evaluations by detecting groups with high agreement in ranking partitioning outcomes; and (c) develops a stable and robust scoring scheme by weighting index values computed across these identified embedding spaces. Experiments show that this new framework aligns better with external measures, effectively reducing the misguidance from the improper use of internal validation measures in deep clustering evaluation. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.

深度聚类(Deep clustering)借助深度神经网络对复杂高维数据执行聚类划分。其流程为先将数据投影至低维嵌入空间,再执行聚类分区,这一过程带来了独有的评估挑战。专为低维空间设计的传统聚类验证指标,应用于深度聚类场景时存在诸多问题,具体原因可归结为两点:(a) 直接应用于高维输入数据时会遭遇维度灾难(curse of dimensionality);(b) 当用于不同嵌入空间生成的嵌入数据时,由于训练流程与模型参数设置存在差异,聚类结果的对比将不再可靠。本文针对深度学习框架下聚类评估中这些尚未解决且常被忽视的难题展开研究。我们提出了一套针对内部聚类验证指标的系统化评估框架,其核心内容包括:(a) 从理论层面阐释传统指标在应用于输入数据,或是跨不同嵌入空间结合聚类分区结果进行评估时失效的根本原因;(b) 通过检测聚类分区结果排序一致性较高的分组,筛选出可支撑可靠评估的嵌入空间;(c) 通过对上述筛选出的嵌入空间所计算得到的指标值进行加权,构建出一套稳定且鲁棒的评分方案。实验结果表明,该全新框架与外部评估指标的对齐效果更优,可有效减少深度聚类评估中因不当使用内部验证指标所带来的误导。本文的补充材料可在线获取,其中包含了可用于复现本研究成果的标准化材料说明。

创建时间:
2025-12-05
二维码
社区交流群
二维码
科研交流群
商业服务