遇见数据集

VFUSE

收藏
DataCite Commons2025-06-01 更新2024-07-27 收录
官方服务:

资源简介:

<br>FUSE is a reproducible, internet-scale corpus, and contains 249,376 unique spreadsheets that were extracted from over 26.83 billion pages. We applied SpreadCluster to the FUSE and manually validated 200 groups that were randomly selected from the clustering result. Based on the validated result, we built the VFUSE corpus, containing 188 evolution groups and 1,143 spreadsheets.VFUSE is published associated with our MSR 2017 paper in May 2017. <br>Liang Xu, Wensheng Dou, Chushu Gao, Jie Wang, Jun Wei, Hua Zhong, Tao Huang. SpreadCluster: Recovering Versioned Spreadsheets through Similarity-Based Clustering. In <i>Proceedings of the 14th International Conference on Mining Software Repositories</i> (<b><i>MSR 2017</i></b>), May 2017.<br>

FUSE是一个可复现的互联网规模语料库,包含249,376个唯一电子表格,这些电子表格源自超过26.83亿个网页页面。我们将SpreadCluster算法应用于FUSE语料库,并从聚类结果中随机选取200个组开展人工验证。基于该验证结果,我们构建了VFUSE语料库,其包含188个演化组与1,143个电子表格。VFUSE于2017年5月随我们的MSR 2017论文一同发布。 徐亮、窦文生、高楚舒、王洁、魏俊、钟华、黄涛。《SpreadCluster:基于相似性聚类的版本化电子表格恢复方法》,收录于《第14届国际软件仓库挖掘会议(International Conference on Mining Software Repositories,MSR 2017)论文集》,2017年5月。

提供机构:
figshare
创建时间:
2017-03-29
搜集汇总
数据集介绍
VFUSE 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务