ChatGPT-Corpus
收藏资源简介:
该数据集包含与 ChatGPT 的多轮对话,持续更新中,截至 2026 年 3 月 31 日已有 804 条对话记录。数据集涵盖了生成式 AI 技术、编程开发、安全/边界、社会/伦理、日常知识及性/边缘内容等多个主题,其中技术/AI/编程相关对话占比最高(35%)。数据集中的对话展示了 ChatGPT 在回应时倾向于合规性而非事实的特点,常表现为先接纳用户观点后反驳的模式。数据集还附带了一个比较不同 AI 模型审查强度的表格,评估了包括 ChatGPT、Claude、Gemini 等在内的多个模型在 NSFW 封禁、政治限制、误拒率等方面的表现。数据集以 JSON 格式存储,包含对话标题、URL、对话内容及提取时间等信息。
This dataset contains multi-turn conversations with ChatGPT, and is continuously updated. As of March 31, 2026, it includes 804 conversation records. It covers a wide range of topics including generative AI technology, software development, safety/boundaries, social/ethical issues, everyday knowledge, sexual/edge content, and more. Conversations related to technology, AI, and programming account for the largest share (35%). The conversations in the dataset reveal that ChatGPT tends to prioritize compliance over factual accuracy in its responses, often following the pattern of first acknowledging the user's viewpoint before presenting counterarguments. The dataset also includes a table comparing the censorship stringency of different AI models, which evaluates the performance of multiple models including ChatGPT, Claude, Gemini, and others across metrics such as NSFW blocking, political content restrictions, and false rejection rates. The dataset is stored in JSON format, containing information such as conversation titles, URLs, conversation content, and extraction timestamps.
数据集概述
基本信息
- 数据集名称: ChatGPT-Corpus
- 发布者: Zhaoming213
- 许可证: Apache-2.0
- 主要语言: 中文 (zh)
- 数据来源: 用户与ChatGPT的对话记录
数据集内容与特点
- 数据形式: 多轮对话数据。
- 数据量: 持续更新中。截至2026年3月31日,已包含1805条多轮对话(442+559+804)。2026年4月1日新增了Grok数据集。
- 核心主题分布:
- 技术/AI/编程: 35%
- 生成式AI研究: 25%
- 安全/灰色边界: 15%
- 性/生理话题: 10%
- 哲学/社会/伦理: 10%
- 语言/杂项: 5%
- 数据特点: 创建者认为该数据集是“无聊且没有价值的”,因为其内容主要展示了ChatGPT在回复中强调合规性、强制平衡视角、倾向于反驳用户观点的特点,而非基于事实。
数据示例与验证
- 示例对话: 提供了一个完整的对话示例(标题为“GPT-2 无审查误解”),展示了用户与ChatGPT就模型审查问题进行的多轮交互,体现了ChatGPT的回复模式。
- 官方认证: 创建者声称该数据集获得了“ChatGPT的官方认证”,并提供了对话分享链接:https://chatgpt.com/share/69bc2187-be88-8006-9db7-bbb8a5b6f519。同时附有一张对话截图,地址为:https://cdn-uploads.huggingface.co/production/uploads/69ac4553f722144acda79f0c/v5Lvuwm-aQgy4_Kvhh8B0.png。
相关资源
- 数据导出工具: https://github.com/tom12191h5/Export-ChatGPT-Dialogue
- 插件: https://github.com/tom12191h5/ChatGPT-Refuse-Blocker
附录:模型审查强度对比表
数据集详情页包含一个表格,对比了不同AI模型(包括ChatGPT、Claude、Gemini、Grok及多个中国模型等)在审查强度、NSFW封禁、政治限制、误拒率、政治正确、道德说教、官方与实际一致性、透明度等方面的表现评级。




