遇见数据集

Yashodhar29/10k-jokes-dataset

收藏
Hugging Face2025-12-12 更新2025-12-20 收录
官方服务:

资源简介:

该数据集包含从原始reddit_jokes.json数据集(195K笑话)中抽取的前10,000个笑话。数据集文件包括原始文件和提取的子集文件,可能包含title、body、score、created_at等字段。适用于幽默生成模型、NLP实验、LLM微调和笑话分类任务等用途。

This dataset contains 10,000 jokes sampled from the original reddit_jokes.json dataset (195K jokes). Only the first 10,000 jokes were selected. The dataset files include the original file and an extracted subset file, with possible fields such as title, body, score, created_at. Useful for humor generation models, NLP experiments, LLM fine-tuning, and joke classification tasks.

提供机构:
Yashodhar29
二维码
社区交流群
二维码
科研交流群
商业服务