DollyTails-12K
收藏资源简介:
DollyTails-12K数据集设计用于遵循指令的任务,采用System 2(类似于O1)的思维范式。数据集中的提示来源于databricks/databricks-dolly-15k,并由GPT-4o进行思考和答案的注释。经过仔细的过滤和筛选,最终数据集包含12K个问答对。每个任务平均有4.93个推理步骤,最多不超过7个步骤,以避免训练过程中不必要的开销。该数据集可用于对大型语言模型(LLM)进行监督微调(SFT),以获得具有System 2类似推理范式的模型。
DollyTails-12K dataset is designed for instruction-following tasks, adopting the System 2 thinking paradigm similar to that of O1. The prompts in the dataset are sourced from databricks/databricks-dolly-15k, and annotated with reasoning steps and final answers generated by GPT-4o. After rigorous filtering and screening, the finalized dataset contains 12K question-answer pairs. Each task has an average of 4.93 reasoning steps, with a maximum of 7 steps, to avoid unnecessary overhead during model training. This dataset can be utilized for supervised fine-tuning (SFT) of Large Language Models (LLMs) to develop models equipped with System 2-like reasoning paradigms.




