Humicroedit
收藏资源简介:
Humicroedit数据集是由罗切斯特大学计算机科学系和微软研究AI共同创建,旨在研究计算幽默。该数据集包含15,095条经过编辑的新闻标题,这些标题通过简单的替换编辑变得幽默。数据集内容丰富,每条标题都配有五个幽默评分,来源于精心筛选的编辑和评委。创建过程中,研究人员从Reddit收集新闻标题,并通过Amazon Mechanical Turk平台招募专家进行编辑和评分。该数据集适用于多种幽默研究任务,如幽默生成和个性化幽默推荐,旨在解决计算幽默领域的挑战,如幽默检测和生成。
The Humicroedit Dataset was co-created by the Department of Computer Science at the University of Rochester and Microsoft Research AI, with the goal of advancing computational humor research. It comprises 15,095 edited news headlines that have been rendered humorous through simple substitution edits. The dataset is comprehensive in scope, with each headline paired with five humor scores sourced from carefully vetted editors and judges. During the dataset development process, researchers collected news headlines from Reddit, and recruited experts via the Amazon Mechanical Turk platform to perform headline editing and scoring tasks. This dataset is applicable to a variety of humor-related research tasks, such as humor generation and personalized humor recommendation, and aims to address key challenges in the field of computational humor, including humor detection and humor generation.

- 1"President Vows to Cut <Taxes> Hair": Dataset and Analysis of Creative Text Editing for Humorous Headlines罗切斯特大学计算机科学系 · 2019年



