遇见数据集

jrodrigues0801/WildChat

收藏
Hugging Face2026-05-24 更新2026-05-31 收录
官方服务:

资源简介:

WildChat是一个包含65万次人类用户与ChatGPT之间对话的数据集。该数据集通过为用户提供免费访问OpenAI的GPT-3.5和GPT-4来收集,涵盖了广泛的用户-聊天机器人交互类型,这些类型在其他指令微调数据集中未充分覆盖,例如模糊用户请求、代码切换、话题切换、政治讨论等。WildChat既可作为指令微调的数据集,也可作为研究用户行为的宝贵资源。注意:此版本的数据集仅包含非毒性的用户输入和ChatGPT响应。数据集包含多语言内容,检测到66种语言,并已通过Microsoft Presidio和作者手写规则进行去标识化处理。数据字段包括对话ID、模型类型、时间戳、对话内容列表(含角色、内容、语言、毒性标记和去标识化标记)、对话轮次、主要语言、OpenAI审核结果、Detoxify审核结果、毒性标记和去标识化标记。数据集已更新移除有毒对话和个人可识别信息(PII)。

WildChat is a dataset containing 650,000 conversations between human users and ChatGPT. This dataset was collected by providing free access to OpenAI's GPT-3.5 and GPT-4 for users, covering a wide range of user-chatbot interaction types that are underrepresented in other instruction-tuning datasets, such as ambiguous user requests, code-switching, topic shifting, political discussions, and so on. WildChat can serve both as a dataset for instruction tuning and a valuable resource for researching user behavior. Note: This version of the dataset only contains non-toxic user inputs and ChatGPT responses. The dataset includes multilingual content, with 66 languages detected, and has been de-identified using Microsoft Presidio and author-written rules. The data fields include conversation ID, model type, timestamp, list of conversation turns (containing role, content, language, toxicity label and de-identification label), number of conversation turns, primary language, OpenAI moderation results, Detoxify moderation results, toxicity label and de-identification label. The dataset has been updated to remove toxic conversations and personally identifiable information (PII).

提供机构:
jrodrigues0801
二维码
社区交流群
二维码
科研交流群
商业服务