遇见数据集

Dataset for: Some Remarks on the R<sup>2</sup> for Clustering

收藏
DataCite Commons2020-08-30 更新2024-08-17 收录
官方服务:

资源简介:

A common descriptive statistic in cluster analysis is the $R^2$ that measures the overall proportion of variance explained by the cluster means. This note highlights properties of the $R^2$ for clustering. In particular, we show that generally the $R^2$ can be artificially inflated by linearly transforming the data by ``stretching'' and by projecting. Also, the $R^2$ for clustering will often be a poor measure of clustering quality in high-dimensional settings. We also investigate the $R^2$ for clustering for misspecified models. Several simulation illustrations are provided highlighting weaknesses in the clustering $R^2$, especially in high-dimensional settings. A functional data example is given showing how that $R^2$ for clustering can vary dramatically depending on how the curves are estimated.

提供机构:
Wiley
创建时间:
2018-04-11
二维码
社区交流群
二维码
科研交流群
商业服务