The contributors have provided two related datasets, which together constitute the FUSE spreadsheet corpus2. + A Web Analysis dataset of 2,127,284 URLs that return spreadsheet content, along with the
FUSE is a reproducible, internet-scale corpus, and contains 249,376 unique spreadsheets that were extracted from over 26.83 billion pages. We applied SpreadCluster to the FUSE and manually validat