遇见数据集

VFUSE

收藏
Figshare2017-03-29 更新2026-04-29 收录
官方服务:

资源简介:

FUSE is a reproducible, internet-scale corpus, and contains 249,376 unique spreadsheets that were extracted from over 26.83 billion pages. We applied SpreadCluster to the FUSE and manually validated 200 groups that were randomly selected from the clustering result. Based on the validated result, we built the VFUSE corpus, containing 188 evolution groups and 1,143 spreadsheets.VFUSE is published associated with our MSR 2017 paper in May 2017. Liang Xu, Wensheng Dou, Chushu Gao, Jie Wang, Jun Wei, Hua Zhong, Tao Huang. SpreadCluster: Recovering Versioned Spreadsheets through Similarity-Based Clustering. In Proceedings of the 14th International Conference on Mining Software Repositories (MSR 2017), May 2017.

创建时间:
2017-03-29
二维码
社区交流群
二维码
科研交流群
商业服务