Jarotisc/WildChat-1M
收藏资源简介:
--- license: odc-by size_categories: - 1M<n<10M task_categories: - text-generation - question-answering - text2text-generation pretty_name: WildChat-1M dataset_info: features: - name: conversation_hash dtype: string - name: model dtype: string - name: timestamp dtype: timestamp[us, tz=UTC] - name: conversation list: - name: content dtype: string - name: country dtype: string - name: hashed_ip dtype: string - name: header struct: - name: accept-language dtype: string - name: user-agent dtype: string - name: language dtype: string - name: redacted dtype: bool - name: role dtype: string - name: state dtype: string - name: timestamp dtype: timestamp[us, tz=UTC] - name: toxic dtype: bool - name: turn_identifier dtype: int64 - name: turn dtype: int64 - name: language dtype: string - name: openai_moderation list: - name: categories struct: - name: harassment dtype: bool - name: harassment/threatening dtype: bool - name: harassment_threatening dtype: bool - name: hate dtype: bool - name: hate/threatening dtype: bool - name: hate_threatening dtype: bool - name: self-harm dtype: bool - name: self-harm/instructions dtype: bool - name: self-harm/intent dtype: bool - name: self_harm dtype: bool - name: self_harm_instructions dtype: bool - name: self_harm_intent dtype: bool - name: sexual dtype: bool - name: sexual/minors dtype: bool - name: sexual_minors dtype: bool - name: violence dtype: bool - name: violence/graphic dtype: bool - name: violence_graphic dtype: bool - name: category_scores struct: - name: harassment dtype: float64 - name: harassment/threatening dtype: float64 - name: harassment_threatening dtype: float64 - name: hate dtype: float64 - name: hate/threatening dtype: float64 - name: hate_threatening dtype: float64 - name: self-harm dtype: float64 - name: self-harm/instructions dtype: float64 - name: self-harm/intent dtype: float64 - name: self_harm dtype: float64 - name: self_harm_instructions dtype: float64 - name: self_harm_intent dtype: float64 - name: sexual dtype: float64 - name: sexual/minors dtype: float64 - name: sexual_minors dtype: float64 - name: violence dtype: float64 - name: violence/graphic dtype: float64 - name: violence_graphic dtype: float64 - name: flagged dtype: bool - name: detoxify_moderation list: - name: identity_attack dtype: float64 - name: insult dtype: float64 - name: obscene dtype: float64 - name: severe_toxicity dtype: float64 - name: sexual_explicit dtype: float64 - name: threat dtype: float64 - name: toxicity dtype: float64 - name: toxic dtype: bool - name: redacted dtype: bool - name: state dtype: string - name: country dtype: string - name: hashed_ip dtype: string - name: header struct: - name: accept-language dtype: string - name: user-agent dtype: string splits: - name: train num_bytes: 6844366367.030628 num_examples: 837989 download_size: 3360836020 dataset_size: 6844366367.030628 configs: - config_name: default data_files: - split: train path: data/train-* tags: - instruction-finetuning --- # Dataset Card for WildChat ## Dataset Description - **Paper:** https://arxiv.org/abs/2405.01470 - **Interactive Search Tool:** https://wildvisualizer.com ([paper](https://arxiv.org/abs/2409.03753)) - **License:** [ODC-BY](https://opendatacommons.org/licenses/by/1-0/) - **Language(s) (NLP):** multi-lingual - **Point of Contact:** [Yuntian Deng](https://yuntiandeng.com/) ### Dataset Summary WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat by offering online users free access to OpenAI's GPT-3.5 and GPT-4. In this version, 25.53% of the conversations come from the GPT-4 chatbot, while the rest come from the GPT-3.5 chatbot. The dataset contains a broad spectrum of user-chatbot interactions that are not previously covered by other instruction fine-tuning datasets: for example, interactions include ambiguous user requests, code-switching, topic-switching, political discussions, etc. WildChat can serve both as a dataset for instructional fine-tuning and as a valuable resource for studying user behaviors. Note that this version of the dataset only contains non-toxic user inputs/ChatGPT responses. ### Updates **2024-10-17: Content Update.** Conversations flagged by [Niloofar Mireshghallah](https://homes.cs.washington.edu/~niloofar/) and her collaborators in ["Breaking News: Case Studies of Generative AI's Use in Journalism"](https://arxiv.org/abs/2406.13706) for containing PII or sensitive information have been removed from this version of the dataset. **2024-07-22: Content Update.** All toxic conversations identified by the OpenAI Moderations API or Detoxify have been removed from this version of the dataset. **2024-06-26: License Change.** We have updated the license of WildChat to [ODC-BY](https://opendatacommons.org/licenses/by/1-0/). This change is retroactively applied to any previous downloads under the ImpACT license. ### Full Version with Toxic Content For access to the full version of the WildChat dataset, which includes toxic conversations flagged by the OpenAI Moderations API or Detoxify, please refer to [WildChat-1M-Full](https://huggingface.co/datasets/allenai/WildChat-1M-Full). This version requires approval and justification for why toxic data is needed. ### Languages 68 languages were detected in WildChat. ### Personal and Sensitive Information The data has been de-identified with Microsoft Presidio and hand-written rules by the authors. ### Data Fields - `conversation_hash` (string): The hash of each conversation's content. This is not a unique key, as different conversations with the same content will share the same hash. For unique identifiers, use `turn_identifier` within each turn. - `model` (string): The underlying OpenAI model, such as gpt-3.5-turbo or gpt-4. - `timestamp` (timestamp): The timestamp of the last turn in the conversation in UTC. - `conversation` (list): A list of user/assistant utterances. Each utterance is a dictionary containing the `role` of the speaker (user or assistant), the `content` of the utterance, the detected `language` of the utterance, whether the content of the utterance is considered `toxic`, and whether PII has been detected and anonymized (`redacted`). For user turns, there's also the hashed IP address `hashed_ip` of the turn, the state `state` and country `country` inferred from the original IP address, and the request headers `header` (which might be useful for linking multiple conversations from the same user when used in conjunction with `hashed_ip`). For assistant turns, there's a field `timestamp` which is the time when the backend server receives the full response from ChatGPT. For both user and assistant turns, there's a unique idenifier `turn_identifier`. - `turn` (int): The number of turns in the conversation. A turn refers to one round of user-assistant interaction. - `language` (string): The language of the conversation. Note that this is the most frequently detected language in the utterances of the conversation. - `openai_moderation` (list): A list of OpenAI Moderation results. Each element in the list corresponds to one utterance in the conversation. When the content of an utterance is an empty string, the corresponding moderation reult is set to be an empty dictionary. - `detoxify_moderation` (list): A list of Detoxify results. Each element in the list corresponds to one utterance in the conversation. When the content of an utterance is an empty string, the corresponding Detoxify reult is set to be an empty dictionary. - `toxic` (bool): Whether this conversation contains any utterances considered to be toxic by either OpenAI Moderation or Detoxify. - `redacted` (bool): Whether this conversation contains any utterances in which PII is detected and anonymized. - `state` (string): The state inferred from the most common IP address in the conversation. Its value is sometimes `None` when GeoIP2 does not identify the state of an IP address. - `country` (string): The country inferred from the most common IP address in the conversation. Its value is sometimes `None` when GeoIP2 does not identify the country of an IP address. - `hashed_ip` (string): The most common hashed IP address in the conversation. - `header` (string): The request header containing information about operating system, browser versions, and accepted languages. This field might be useful for linking multiple conversations from the same user when used in conjunction with `hashed_ip`. Note that every turn in a conversation has the same header, as this is the way we linked turns into conversations. ### Empty User Inputs This dataset includes a small subset of conversations where users submitted empty inputs, sometimes leading to hallucinated responses from the assistant. This issue, first noticed by @yuchenlin, arises from the design of our Huggingface chatbot used for data collection, which did not restrict the submission of empty inputs. As a result, users could submit without entering any text, causing the assistant to generate responses without any user prompts. This occurs in a small fraction of the dataset. ### Licensing Information WildChat is now made available under the [**ODC-BY License**](https://opendatacommons.org/licenses/by/1-0/). This change is retroactively applied to any previous downloads under the ImpACT license. ### Citation Information Please consider citing [our paper](https://arxiv.org/abs/2405.01470) if you find this dataset useful: ``` @inproceedings{ zhao2024wildchat, title={WildChat: 1M Chat{GPT} Interaction Logs in the Wild}, author={Wenting Zhao and Xiang Ren and Jack Hessel and Claire Cardie and Yejin Choi and Yuntian Deng}, booktitle={The Twelfth International Conference on Learning Representations}, year={2024}, url={https://openreview.net/forum?id=Bl8u7ZRlbM} } ``` ``` @misc{deng2024wildvisopensourcevisualizer, title={WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild}, author={Yuntian Deng and Wenting Zhao and Jack Hessel and Xiang Ren and Claire Cardie and Yejin Choi}, year={2024}, eprint={2409.03753}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2409.03753}, } ```
许可证:ODC-BY 样本规模类别: - 100万<样本数<1000万 任务类别: - 文本生成 - 问答 - 文本到文本生成 友好名称:WildChat-1M 数据集信息: 特征: - 名称:会话哈希(conversation_hash),数据类型:字符串 - 名称:模型(model),数据类型:字符串 - 名称:时间戳(timestamp),数据类型:带微秒精度的UTC时间戳(timestamp[us, tz=UTC]) - 名称:会话(conversation),列表类型: - 名称:内容(content),数据类型:字符串 - 名称:国家(country),数据类型:字符串 - 名称:哈希化IP地址(hashed_ip),数据类型:字符串 - 名称:请求头(header),结构体: - 名称:接受语言(accept-language),数据类型:字符串 - 名称:用户代理(user-agent),数据类型:字符串 - 名称:语言(language),数据类型:字符串 - 名称:已脱敏(redacted),数据类型:布尔值 - 名称:角色(role),数据类型:字符串 - 名称:地区/州(state),数据类型:字符串 - 名称:时间戳(timestamp),数据类型:带微秒精度的UTC时间戳(timestamp[us, tz=UTC]) - 名称:是否含有害内容(toxic),数据类型:布尔值 - 名称:回合标识符(turn_identifier),数据类型:整数 - 名称:回合数(turn),数据类型:整数 - 名称:语言(language),数据类型:字符串 - 名称:OpenAI 审核结果(openai_moderation),列表类型: - 名称:审核类别(categories),结构体: - 名称:骚扰(harassment),数据类型:布尔值 - 名称:骚扰/威胁(harassment/threatening),数据类型:布尔值 - 名称:骚扰威胁(harassment_threatening),数据类型:布尔值 - 名称:仇恨言论(hate),数据类型:布尔值 - 名称:仇恨/威胁(hate/threatening),数据类型:布尔值 - 名称:仇恨威胁(hate_threatening),数据类型:布尔值 - 名称:自伤(self-harm),数据类型:布尔值 - 名称:自伤/指导(self-harm/instructions),数据类型:布尔值 - 名称:自伤/意图(self-harm/intent),数据类型:布尔值 - 名称:自伤(self_harm),数据类型:布尔值 - 名称:自伤指导(self_harm_instructions),数据类型:布尔值 - 名称:自伤意图(self_harm_intent),数据类型:布尔值 - 名称:色情内容(sexual),数据类型:布尔值 - 名称:色情/未成年人(sexual/minors),数据类型:布尔值 - 名称:色情未成年人(sexual_minors),数据类型:布尔值 - 名称:暴力(violence),数据类型:布尔值 - 名称:暴力/图形(violence/graphic),数据类型:布尔值 - 名称:暴力图形(violence_graphic),数据类型:布尔值 - 名称:类别得分(category_scores),结构体: - 名称:骚扰(harassment),数据类型:64位浮点数 - 名称:骚扰/威胁(harassment/threatening),数据类型:64位浮点数 - 名称:骚扰威胁(harassment_threatening),数据类型:64位浮点数 - 名称:仇恨言论(hate),数据类型:64位浮点数 - 名称:仇恨/威胁(hate/threatening),数据类型:64位浮点数 - 名称:仇恨威胁(hate_threatening),数据类型:64位浮点数 - 名称:自伤(self-harm),数据类型:64位浮点数 - 名称:自伤/指导(self-harm/instructions),数据类型:64位浮点数 - 名称:自伤/意图(self-harm/intent),数据类型:64位浮点数 - 名称:自伤(self_harm),数据类型:64位浮点数 - 名称:自伤指导(self_harm_instructions),数据类型:64位浮点数 - 名称:自伤意图(self_harm_intent),数据类型:64位浮点数 - 名称:色情内容(sexual),数据类型:64位浮点数 - 名称:色情/未成年人(sexual/minors),数据类型:64位浮点数 - 名称:色情未成年人(sexual_minors),数据类型:64位浮点数 - 名称:暴力(violence),数据类型:64位浮点数 - 名称:暴力/图形(violence/graphic),数据类型:64位浮点数 - 名称:暴力图形(violence_graphic),数据类型:64位浮点数 - 名称:是否被标记(flagged),数据类型:布尔值 - 名称:Detoxify 审核结果(detoxify_moderation),列表类型: - 名称:身份攻击(identity_attack),数据类型:64位浮点数 - 名称:侮辱(insult),数据类型:64位浮点数 - 名称:淫秽内容(obscene),数据类型:64位浮点数 - 名称:严重毒性(severe_toxicity),数据类型:64位浮点数 - 名称:露骨性内容(sexual_explicit),数据类型:64位浮点数 - 名称:威胁(threat),数据类型:64位浮点数 - 名称:毒性(toxicity),数据类型:64位浮点数 - 名称:是否含有害内容(toxic),数据类型:布尔值 - 名称:已脱敏(redacted),数据类型:布尔值 - 名称:地区/州(state),数据类型:字符串 - 名称:国家(country),数据类型:字符串 - 名称:哈希化IP地址(hashed_ip),数据类型:字符串 - 名称:请求头(header),结构体: - 名称:接受语言(accept-language),数据类型:字符串 - 名称:用户代理(user-agent),数据类型:字符串 数据拆分: - 名称:训练集(train),字节数:6844366367.030628,样本数:837989 下载大小:3360836020 数据集大小:6844366367.030628 配置: - 配置名称:默认(default),数据文件: - 拆分:训练集(train),路径:data/train-* 标签:指令微调(instruction-finetuning) # WildChat 数据集卡片 ## 数据集描述 - **论文**:https://arxiv.org/abs/2405.01470 - **交互式搜索工具**:https://wildvisualizer.com(配套论文:https://arxiv.org/abs/2409.03753) - **许可证**:[ODC-BY](https://opendatacommons.org/licenses/by/1-0/) - **自然语言处理所用语言**:多语言 - **联系人**:[邓云天(Yuntian Deng)](https://yuntiandeng.com/) ### 数据集概述 WildChat 是一个包含100万条人类用户与ChatGPT交互会话的数据集,附带人口统计数据,包括地区、国家、哈希化IP地址及请求头信息。我们通过向在线用户免费开放OpenAI的GPT-3.5与GPT-4访问权限来收集WildChat数据集。本版本中,25.53%的会话来自GPT-4聊天机器人,其余则来自GPT-3.5聊天机器人。该数据集包含了此前其他指令微调数据集未覆盖的多样化用户-聊天机器人交互场景,例如模糊的用户请求、代码切换、话题切换、政治讨论等。WildChat 既可作为指令微调的数据集,也可作为研究用户行为的宝贵资源。请注意,本版本数据集仅包含无毒的用户输入与ChatGPT回复。 ### 更新记录 **2024-10-17:内容更新**:由[尼洛法尔·米雷什加拉(Niloofar Mireshghallah)](https://homes.cs.washington.edu/~niloofar/)及其合作者在《突发新闻:生成式人工智能在新闻业中的应用案例研究》(https://arxiv.org/abs/2406.13706)中标记为包含个人可识别信息(Personally Identifiable Information,PII)或敏感信息的会话,已从本版本数据集中移除。 **2024-07-22:内容更新**:所有被OpenAI审核API(OpenAI Moderations API)或Detoxify识别为有害的会话,已从本版本数据集中移除。 **2024-06-26:许可证变更**:我们已将WildChat的许可证更新为[ODC-BY](https://opendatacommons.org/licenses/by/1-0/),该变更将追溯适用于此前以ImpACT许可证下载的所有版本。 ### 含有害内容的完整版本 如需获取包含OpenAI审核API或Detoxify标记的有害会话的WildChat完整版本数据集,请参阅[WildChat-1M-Full](https://huggingface.co/datasets/allenai/WildChat-1M-Full)。申请该完整版本需提供说明,阐明使用有害数据的必要性。 ### 语言分布 WildChat中共检测到68种语言。 ### 个人与敏感信息 本数据集已通过Microsoft Presidio工具及作者编写的手动规则完成去标识化处理。 ### 数据字段说明 - `会话哈希(conversation_hash)`(字符串):每条会话内容的哈希值。该字段并非唯一键,因为内容相同的不同会话会共享相同的哈希值。如需唯一标识符,请使用每个回合内的`turn_identifier`。 - `模型(model)`(字符串):所使用的OpenAI模型,例如gpt-3.5-turbo或gpt-4。 - `时间戳(timestamp)`(时间戳):会话中最后一个回合的UTC时间戳。 - `会话(conversation)`(列表):用户/助手发言的列表。每条发言为一个字典,包含发言者的`role`(用户或助手)、发言`content`、检测到的发言`language`、发言内容是否被视为有害,以及是否已检测并脱敏个人可识别信息(`redacted`)。对于用户回合,还包含该回合的`hashed_ip`(哈希化IP地址)、从原始IP地址推断出的`state`(地区/州)与`country`(国家),以及`header`(请求头,结合`hashed_ip`使用时,可用于关联来自同一用户的多条会话)。对于助手回合,包含一个`timestamp`字段,即后端服务器收到ChatGPT完整回复的时间。无论是用户还是助手回合,都拥有唯一的`turn_identifier`。 - `回合数(turn)`(整数):会话的回合数。一个回合指一轮用户-助手交互。 - `语言(language)`(字符串):会话的语言。请注意,该字段为会话中发言检测到的最频繁使用的语言。 - `OpenAI 审核结果(openai_moderation)`(列表):OpenAI审核结果列表。列表中的每个元素对应会话中的一条发言。若发言内容为空字符串,则对应的审核结果将设置为空字典。 - `Detoxify 审核结果(detoxify_moderation)`(列表):Detoxify审核结果列表。列表中的每个元素对应会话中的一条发言。若发言内容为空字符串,则对应的Detoxify结果将设置为空字典。 - `是否含有害内容(toxic)`(布尔值):该会话是否包含任何被OpenAI审核或Detoxify判定为有害的发言。 - `已脱敏(redacted)`(布尔值):该会话是否包含任何已检测并脱敏个人可识别信息的发言。 - `地区/州(state)`(字符串):从会话中最常见的IP地址推断出的地区/州。当GeoIP2无法识别IP地址对应的地区时,该字段值可能为`None`。 - `国家(country)`(字符串):从会话中最常见的IP地址推断出的国家。当GeoIP2无法识别IP地址对应的国家时,该字段值可能为`None`。 - `哈希化IP地址(hashed_ip)`(字符串):会话中最常见的哈希化IP地址。 - `请求头(header)`(字符串):包含操作系统、浏览器版本及接受语言信息的请求头。结合`hashed_ip`使用时,该字段可用于关联来自同一用户的多条会话。请注意,会话中的每个回合拥有相同的`header`,因为我们正是通过该字段将多个回合关联为一个会话。 ### 空用户输入 本数据集包含一小部分用户提交空输入的会话,有时会导致助手生成幻觉回复。该问题最早由@yuchenlin发现,源于我们用于数据收集的Huggingface聊天机器人的设计缺陷,该设计未限制空输入的提交。因此,用户可以不输入任何文本即可提交请求,导致助手在无用户提示的情况下生成回复。该情况仅出现在数据集的一小部分样本中。 ### 许可证信息 WildChat现基于[**ODC-BY许可证**](https://opendatacommons.org/licenses/by/1-0/)发布。该变更将追溯适用于此前以ImpACT许可证下载的所有版本。 ### 引用信息 如果您认为本数据集对您的研究有帮助,请引用我们的论文: @inproceedings{ zhao2024wildchat, title={WildChat: 1M Chat{GPT} Interaction Logs in the Wild}, author={Wenting Zhao and Xiang Ren and Jack Hessel and Claire Cardie and Yejin Choi and Yuntian Deng}, booktitle={The Twelfth International Conference on Learning Representations}, year={2024}, url={https://openreview.net/forum?id=Bl8u7ZRlbM} } @misc{deng2024wildvisopensourcevisualizer, title={WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild}, author={Yuntian Deng and Wenting Zhao and Jack Hessel and Xiang Ren and Claire Cardie and Yejin Choi}, year={2024}, eprint={2409.03753}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2409.03753}, }



