Ethosoft/nedo-turkish-sft-mixtures
收藏资源简介:
NEDO Turkish SFT Mixtures是一个土耳其语监督微调数据集,专为NEDO Turkish SLM项目设计。该数据集用于在NEDO Turkish 65K标记化网络语料库基础预训练后,微调NEDOQwen风格的土耳其语仅解码器语言模型。数据集包含两个主要文件:tr_sft_clean_20k.jsonl(20,000个示例,推荐使用,为更干净的土耳其语指令子集)和tr_sft_hf_mix.jsonl(104,078个示例,实验性,为更大但更嘈杂的土耳其语指令混合集)。数据格式为JSONL,每条记录包含instruction(用户指令)、input(可选额外上下文)、output(目标助手响应)和source(上游数据源标识)字段。数据集支持多种任务类型,如文本生成、问答、摘要和文本到文本生成,主要用于土耳其语监督微调、小型语言模型对齐实验、指令跟随研究以及土耳其SLM评估和消融研究。数据集基于上游土耳其语指令数据集(如NovusResearch/turkish_instructions、SoAp9035/turkish_instructions等)构建,用户需遵守上游许可条款。
NEDO Turkish SFT Mixtures is a Turkish supervised fine-tuning dataset prepared for the NEDO Turkish SLM project. It is designed to fine-tune NEDOQwen-style Turkish decoder-only language models after base pretraining on the NEDO Turkish 65K tokenized web corpus. The dataset includes two main files: tr_sft_clean_20k.jsonl (20,000 examples, recommended as a cleaner Turkish instruction subset) and tr_sft_hf_mix.jsonl (104,078 examples, experimental, as a larger but noisier Turkish instruction mixture). The data format is JSONL, with each record containing fields: instruction (user instruction), input (optional extra context), output (target assistant response), and source (upstream dataset identifier). It supports multiple task categories such as text generation, question answering, summarization, and text2text generation. The dataset is intended for Turkish supervised fine-tuning, small language model alignment experiments, instruction-following research, and Turkish SLM evaluation and ablation studies. It is derived from upstream Turkish instruction datasets (e.g., NovusResearch/turkish_instructions, SoAp9035/turkish_instructions), and users are responsible for complying with upstream licenses.



