遇见数据集

Data for: kluster: An Efficient Scalable Procedure for Approximating the Number of Clusters in Unsupervised Learning

收藏
Mendeley Data2018-06-19 更新2026-04-09 收录
官方服务:

资源简介:

182 simulated datasets (first set contains small datasets and second set contains large datasets) with different cluster compositions – i.e., different number clusters and separation values – generated using clusterGeneration package in R. Each set of simulation datasets consists of 91 datasets in comma separated values (csv) format (total of 182 csv files) with 3-15 clusters and 0.1 to 0.7 separation values. Separation values can range between (−0.999, 0.999), where a higher separation value indicates cluster structure with more separable clusters. Size of the dataset, number of clusters, and separation value of the clusters in the dataset is printed in file name. size_X_n_Y_sepval_Z.csv: Size of the dataset = X number of clusters in the dataset = Y separation value of the clusters in the dataset = Z

本数据集共包含182个模拟数据集,分为两组:第一组为小数据集,第二组为大数据集,所有数据集均采用不同的聚类组成——即聚类数量与分离度参数各不相同——通过R语言的clusterGeneration包生成。 每组模拟数据集包含91个以逗号分隔值(CSV)格式存储的文件(总计182个CSV文件),涵盖3至15个聚类,分离度参数取值为0.1至0.7。分离度参数的理论取值区间为(-0.999, 0.999),数值越高,代表聚类结构的可分离性越强。 数据集的规模、聚类数量以及聚类分离度参数均会在文件名中体现,命名格式为:size_X_n_Y_sepval_Z.csv,其中X为数据集规模,Y为数据集内的聚类数量,Z为数据集的聚类分离度参数。

创建时间:
2018-06-19
二维码
社区交流群
二维码
科研交流群
商业服务