遇见数据集

On Exploring Data Lakes by Finding Compact, Isolated Clusters

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

These are the research materials that accompany article "On Exploring Data Lakes by Finding Compact, Isolated Clusters", by Patricia Jiménez, Juan C. Roldán, and Rafael Corchuelo. This package includes the following: - "system": it provides the python code required to run and test RóMULO. There is a "launch.cmd" script that launches the experimentation. The implementation of the competitors can be found elsewhere. The implementation of GSPPCA is available from the authors at https://github.com/pamattei/GSPPCA. The implementation of AffinityPropagation, Meanshift, and OPTICS-XI is available from SckitLearn at https://scikit-learn.org/stable/install.html. The implementation of PQC is available from the authors at https://github.com/racaes/PQC. The implementation of DCC is also available from the authors at https://github.com/shahsohil/DCC. - "data-lakes": each subfolder corresponds to a data lake, and each CSV file inside a data-lake corresponds to a dataset. The data lakes in package "clustering.zip" are intended to evaluate the proposal regarding unsupervised quality coefficients (the class attribute is set to zero in all cases). The data lakes in package "classification.zip" are intended to evaluate the proposal regarding supervised quality coefficients (the class attributed is encoded using an enumerated natural number). - "results": it provides the results of evaluating RóMULO and other competitors on the previous data lakes.

本数据集为Patricia Jiménez、Juan C. Roldán及Rafael Corchuelo发表的论文《On Exploring Data Lakes by Finding Compact, Isolated Clusters》配套的研究材料。本数据包包含以下内容: - 系统模块(system):提供运行与测试RóMULO所需的Python代码。内置`launch.cmd`脚本用于启动实验。基线算法的实现可从其他渠道获取:GSPPCA的实现可通过作者仓库获取,地址为https://github.com/pamattei/GSPPCA;亲和力传播(AffinityPropagation)、均值漂移(Meanshift)及OPTICS-XI的实现可从Scikit-learn获取,地址为https://scikit-learn.org/stable/install.html;PQC的实现可通过作者仓库获取,地址为https://github.com/racaes/PQC;DCC的实现同样可通过作者仓库获取,地址为https://github.com/shahsohil/DCC。 - 数据湖(data-lakes):每个子文件夹对应一个数据湖,数据湖内的每个CSV文件对应一个数据集。`clustering.zip`数据包内的数据湖用于评估无监督质量系数相关的研究方案(所有样本的类别属性均设为0);`classification.zip`数据包内的数据湖用于评估有监督质量系数相关的研究方案(类别属性通过枚举自然数进行编码)。 - 结果模块(results):提供RóMULO及其他基线算法在上述数据湖上的评估结果。

创建时间:
2021-10-29
二维码
社区交流群
二维码
科研交流群
商业服务