shuttie/dadjokes
收藏资源简介:
该数据集来源于Kaggle的Reddit Dad Jokes,由Oktay Ozturk生成并进行了修改。只保留了获得5个以上投票的笑话,并将每个笑话分为基础和笑点两部分。数据集格式为CSV,分为训练集(52000个样本)和测试集(1400个样本),可用于笑话预测任务。
This dataset is sourced from Kaggle's Reddit Dad Jokes, and was generated and modified by Oktay Ozturk. Only jokes that received over 5 upvotes were retained, and each joke is split into two parts: the setup and the punchline. The dataset is in CSV format, divided into a training set with 52,000 samples and a test set with 1,400 samples, which can be used for joke prediction tasks.
Dad Jokes 数据集
概述
该数据集源自 Kaggle Reddit Dad Jokes,由 Oktay Ozturk 创建,并进行了以下修改:
- 仅包含获得 5 票以上的笑话,以避免低票笑话的质量问题。
- 通过一系列启发式方法,将每个笑话分为基础部分和笑点部分。
格式
数据集以 CSV 格式提供,并分为训练集和测试集:
- 训练集:52000 个样本
- 测试集:1400 个样本
示例数据
csv "question","response" "I asked my priest how he gets holy water","He said it’s just regular water, he just boils the hell out of it" "Life Hack: If you play My Chemical Romance loud enough in your yard","your grass will cut itself" "Why did Mr. Potato Head get pulled over","He was baked" "How did the Mexican John Wick taste his Burrito","He took Juan Lick"
用途
该数据集可用于基于基础/笑点分割的笑话预测任务,适用于任何大型语言模型(LLM)。
许可证
Apache 2.0。




