COIG-Writer
收藏资源简介:
该数据集收集了大量高质量的中文创意写作作品和其它文本类型(如科普文章),每个作品都附有详细的“Query”(提示)和“Thought”(思考过程)。该数据集旨在解决机器生成文本中常见的“AI风味”问题,如逻辑不一致、缺乏个性、分析肤浅、语言过于复杂或叙事发展薄弱等。主要目标是提供资源,帮助训练语言模型生成内容流畅、具有深度连贯性、个性、洞察力和复杂叙事结构的文本,更接近人类创作的作品。数据集涵盖了大约50个子领域的中文创意写作和其它文本生成任务。所有文本均为简体中文(zh-CN)。每个数据实例包括以下组件:`query_type`、`query`、`thought`、`answer`、`link`和`score`。
This dataset contains a large corpus of high-quality Chinese creative writing works and other text types (e.g., popular science articles), with each piece accompanied by detailed "Query" (prompt) and "Thought" (thinking process). This dataset aims to address common "AI-style" issues in machine-generated text, such as logical inconsistency, lack of personality, superficial analysis, overly complex language, and underdeveloped narrative arcs. Its primary goal is to provide resources to train language models to produce text that is fluent, deeply coherent, personalized, insightful, and structured with complex narrative frameworks, closely resembling human-created works. The dataset covers Chinese creative writing and other text generation tasks across approximately 50 subfields. All texts are in Simplified Chinese (zh-CN). Each data instance includes the following components: `query_type`, `query`, `thought`, `answer`, `link`, and `score`.




