遇见数据集

CAH

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

我们在反人类纸牌的背景下探索幽默-一种派对游戏,玩家使用可能令人反感或政治上不正确的纸牌来完成空白陈述。我们介绍了一个新颖的数据集,其中包括785K个独特的笑话,对其进行分析并提供见解的300,000在线游戏。我们训练了机器学习模型来预测每场比赛的获胜笑话,即使没有任何用户信息,也能达到随机两倍的性能 (20%)。在判断新颖卡片的更艰巨的任务中,我们看到模型的概括能力是中等的。有趣的是,我们发现我们的模型主要集中在punchline card上,上下文几乎没有影响。分析特征的重要性,我们观察到短的,粗糙的,少年的笑点往往会获胜。

We explore humor in the context of Cards Against Humanity—a party game where players use potentially offensive or politically incorrect cards to complete blank statements. We introduce a novel dataset comprising 785K unique jokes and 300,000 online games, which are analyzed to provide research insights. We train machine learning models to predict the winning joke of each game, achieving performance twice that of random guessing (20%) even without any user information. In the more challenging task of judging novel cards, we observe that the model's generalization ability is moderate. Interestingly, we find that our models primarily focus on the punchline cards, with the context exerting little impact. By analyzing feature importance, we observe that short, vulgar, and juvenile punchlines tend to win the games.

提供机构:
OpenDataLab
创建时间:
2022-11-18
搜集汇总
数据集介绍
CAH 数据集图片
背景与挑战
背景概述
CAH数据集基于反人类纸牌游戏,包含785K个独特笑话和30万场在线游戏数据,用于训练模型预测获胜笑话并分析幽默特征。该数据集由耶路撒冷希伯来大学于2022年发布,提供相关论文和开源链接。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务