遇见数据集

Long document similarity dataset, Wikipedia excerptions for wine collections

收藏
Zenodo2021-05-26 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

Wine-related articles extracted from Wikipedia. For all articles, the figures and tables have been filtered out, as well as the categories and "see also" sections. The article structure, and particularly the sub-titles and paragraphs are kept in these datasets <strong>Wines</strong> Wikipedia wines dataset consists of 1635 articles from the wine domain. The extracted dataset consists of a non-trivial mixture of articles, including different wine categories, brands, wineries, grape types, and more. The ground-truth recommendations were crafted by a human sommelier, which annotated 92 source articles with ~10 ground-truth recommendations for each sample. Examples for ground-truth expert-based recommendations are Dom Pérignon - Moët &amp; Chandon Pinot Meunier - Chardonnay

提供机构:
Zenodo
创建时间:
2021-05-26
二维码
社区交流群
二维码
科研交流群
商业服务