SYNTHQUESTIONS
收藏资源简介:
SYNTHQUESTIONS数据集是由中国科学技术大学和Metastone Technology合作构建的,包含100万条经过精心设计和合成的用户指令。该数据集通过一种称为“属性接地”的新颖框架生成,该框架结合了自顶向下的属性过程和自底向上的合成过程,以确保生成的指令既多样化又复杂,能够有效提升大型语言模型的理解和推理能力。数据集的创建过程首先收集了大量真实的人类指令,并进行了严格的清洗和去重,然后利用这些指令作为种子,通过先进的语言模型生成多样化的指令。该数据集在多个基准测试中表现出领先性能,显示出其在提升大型语言模型理解和推理能力方面的巨大潜力。
SYNTHQUESTIONS Dataset is co-constructed by the University of Science and Technology of China and Metastone Technology, comprising 1,000,000 meticulously designed and synthesized user instructions. This dataset is generated through a novel framework named "attribute grounding", which integrates top-down attribute processing and bottom-up synthesis procedures to ensure the generated instructions are both diverse and sophisticated, effectively enhancing the comprehension and reasoning capabilities of large language models. The development pipeline of this dataset first collects a large corpus of real human instructions, followed by strict cleaning and deduplication. Subsequently, these collected instructions are utilized as seeds to generate diverse instructions via state-of-the-art language models. This dataset has demonstrated leading performance across multiple benchmark tests, showcasing its significant potential in enhancing the comprehension and reasoning abilities of large language models.
SynthQuestions数据集概述
基本信息
- 数据集名称:SynthQuestions
- 关联论文:《From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding》
数据集特点
- 数据生成方式:合成生成
- 数据规模:数百万条
- 数据特征:多样化且复杂的用户指令
- 特殊属性:带有属性标注的基础信息
当前状态
- 项目处于未完成状态(标注"WILL BE COMPLETED SOON")

- 1From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding中国科学技术大学, Metastone Technology · 2025年



