WildChat - 100万用户与ChatGPT互动日志的多语种数据集
收藏资源简介:
WildChat数据集由康奈尔大学与艾伦人工智能研究所联合创建,旨在填补对话式AI研究中真实用户与聊天机器人互动数据的空白。该数据集包含了100万用户与ChatGPT的对话,超过250万个交互轮次,并包含了时间戳、人口统计数据(如国家、地区和哈希化的IP地址)以及请求头。此外,该数据集还具有多语言特性,提供了比现有数据集更接近真实世界的多轮对话交互。在数据收集过程中,研究团队通过提供ChatGPT的免费访问,获得了用户的明确同意,匿名收集了聊天记录和请求头信息。该数据集为研究人员提供了宝贵的资源,有助于研究和对抗有害的聊天机器人互动,并在微调指令遵循模型方面展示了潜在的应用价值。
The WildChat dataset was jointly created by Cornell University and the Allen Institute for AI, aiming to fill the gap in real-world user-chatbot interaction data for conversational AI research. It contains 1 million conversations between users and ChatGPT, with over 2.5 million interaction turns, along with timestamps, demographic data (including country, region, and hashed IP addresses), and request headers. Additionally, the dataset features multilingual capabilities, providing multi-turn dialogue interactions that are closer to real-world scenarios than existing datasets. During the data collection process, the research team obtained explicit consent from users by providing free access to ChatGPT, and anonymously collected chat records and request header information. This dataset serves as a valuable resource for researchers, facilitating research on and defense against harmful chatbot interactions, and demonstrating potential applications in fine-tuning instruction-following models.




