Ali-Yaser/CodeX-2M-Thinking-ShareGPT
收藏资源简介:
CodeX-2M-Thinking-ShareGPT是一个从Modotte/CodeX-2M-Thinking源数据集转换到ShareGPT格式的数据集。它包含2,189,671行对话数据,采用JSON格式,其中conversations数组包括human和gpt的对话,涉及编码问题和解决方案。该数据集适用于文本生成任务,特别关注代码生成、思维链和指令调优,语言为英语,规模在100万到1000万之间,兼容Axolotl、LLaMA-Factory和Unsloth等工具。
CodeX-2M-Thinking-ShareGPT is a dataset converted from the Modotte/CodeX-2M-Thinking source to ShareGPT format. It contains 2,189,671 rows of conversation data in JSON format, where the conversations array includes dialogues between human and gpt, covering coding problems and solutions. The dataset is suitable for text-generation tasks, particularly focusing on code generation, thinking chains, and instruction tuning. It is in English, with a size between 1M and 10M, and is compatible with tools like Axolotl, LLaMA-Factory, and Unsloth.



