alibayram/Bilge-Turkish-CoT-50K
收藏资源简介:
Bilge Turkish CoT 50K 是一个包含50,000个样本的土耳其语链式思维推理微调数据集。该数据集旨在提升土耳其语大语言模型的逐步推理能力,每个样本设计为让模型先在<think>...</think>块内进行显式推理,然后提供结构化、详细的回答。数据集是bugrabilge/Omni-31B-Turkish-Reasoning-Model模型训练所用249K过滤数据集的50K子集,内容包含学术和解释性风格的长篇土耳其语回答,涵盖科学、历史、文学、哲学、医学等多个领域。数据格式遵循HuggingFace对话标准,为Messages JSONL,每个样本包含system、user、assistant三种角色的对话。数据集通过多阶段规则过滤流程从约30GB原始土耳其语CoT数据生成,具有领域多样性、目标受众多样性和多种解释方法。
Bilge Turkish CoT 50K is a Turkish Chain-of-Thought reasoning fine-tuning dataset containing 50,000 examples. It is designed to enhance the step-by-step reasoning capabilities of Turkish large language models. Each example is structured to encourage the model to first perform explicit reasoning within <think>...</think> blocks, then provide structured and detailed responses. This dataset is a 50K subset of the 249K filtered dataset used to train the bugrabilge/Omni-31B-Turkish-Reasoning-Model. It contains long-form Turkish answers in an academic and explanatory style, covering various fields such as science, history, literature, philosophy, medicine, and more. The data format follows the HuggingFace conversational standard as Messages JSONL, with each sample containing a three-role conversation (system, user, assistant). The dataset is generated through a multi-stage rule-based filtering pipeline from approximately 30GB of raw Turkish CoT data, featuring domain diversity, target audience diversity, and multiple explanation methods.




