Tralalabs/diverse-dataset-it-162x
收藏资源简介:
Diverse Dataset IT 162x是一个通过LMArena的Battle Mode生成的问答数据集,使用了Claude Opus 4.8 Thinking和MiniMax M2.7模型。注意:MiniMax M2.7生成了部分仅含问题的行,因此使用ChatGPT的GPT-5.5 Instant补充了回答。数据集以JSONL格式存储,包含75行,类型为Q&C&A(问题、上下文、答案)。在35-45%的行中,上下文字段被完全省略(而非留空或设为null)。问题设计多样且自然,涵盖多个领域,如推理、编程、科学、数学、创意写作、历史、地理、技术、健康(非医疗诊断)、生产力、语言、琐事、指令遵循、多步骤任务、总结和比较。难度从易到专家不等,问题长度和详细程度各异。上下文可能包含段落、项目符号信息、摘录、伪造文档、代码片段、表格文本转换和场景描述。答案要求高质量、简洁或详细(视情况而定)、基于事实、自然书写。数据集还包括边缘案例,如模糊问题、不完整上下文、指令密集型提示、上下文中的冲突信息以及长推理任务。避免重复措辞或模板,确保语气和结构的多样性。
A Q&C&A dataset generated by models on LMArena, Battle Mode. Models used to generate: Claude Opus 4.8 Thinking and MiniMax M2.7. Note: MiniMax M2.7 had included rows with only questions. So we generated the rows that only had "question" with GPT-5.5 Instant on ChatGPT. The dataset is in JSONL format with exactly 75 rows, following the Q&C&A (Question, Context, Answer) type. Context is optional and omitted entirely in around 35–45% of rows. Questions are diverse and natural, covering domains such as reasoning, coding, science, math, creative writing, history, geography, technology, health (non-medical diagnosis), productivity, language, trivia, instruction following, multi-step tasks, summarization, and comparisons. Difficulty varies from easy to expert, with a mix of short and long questions. Context may include paragraphs, bullet-style information, excerpts, fake documents, code snippets, tables converted to text, and scenario descriptions. Answers are high quality, concise or detailed as appropriate, factually grounded, and naturally written. Edge cases include ambiguous questions, incomplete context, instruction-heavy prompts, conflicting information in context, and long reasoning tasks. The dataset avoids repetitive wording and ensures variety in tone and structure.




