OpenBahasa-CoT
收藏资源简介:
OpenBahasa-CoT是一个高质量的印度尼西亚语数据集,专门设计用于增强大语言模型在训练过程中的推理能力。它侧重于提升结构化思维、逐步推理和跨多种任务的连贯回答生成,所有示例均使用自然印度尼西亚语编写,以确保与真实世界使用场景和语言模式对齐。数据集中的每个样本都经过精心设计,展示清晰且逻辑性强的思维过程,帮助模型学习如何以一致、可解释的方式推理问题,并将问题分解为逻辑步骤。数据集涵盖基础推理、中级推理、高级推理和复杂推理等多个难度级别,适用于训练/预训练、监督微调、指令微调和推理增强等应用场景。它包含user、reasoning和assistant三个必需字段,其中assistant字段整合了推理过程和最终答案。目前,数据集被分为多个子集,并仍在持续扩展和改进中。
OpenBahasa-CoT is a high-quality Indonesian language dataset specifically designed to enhance the reasoning capabilities of large language models during training. It focuses on improving structured thinking, step-by-step reasoning, and coherent answer generation across multiple tasks. All examples in the dataset are written in natural Indonesian, ensuring better alignment with real-world usage scenarios and language patterns. Each sample is carefully designed to demonstrate clear and logical thought processes, enabling models to better understand how to reason through problems in a consistent, interpretable, and explainable manner. The dataset covers multiple difficulty levels, including basic, intermediate, advanced, and complex reasoning tasks, and is suitable for applications such as training/pre-training, supervised fine-tuning, instruction fine-tuning, and reasoning enhancement pipelines. It includes three required fields: user, reasoning, and assistant, where the assistant field integrates the reasoning process and final answer. Currently, the dataset is divided into multiple subsets (configurations) and is continually being expanded and improved.
数据集概述:OpenBahasa-CoT
OpenBahasa-CoT 是一个高质量的印尼语数据集,旨在提升大语言模型在训练过程中的推理能力。该数据集侧重于增强结构化思维、逐步推理以及跨多种任务的连贯响应生成能力。
核心特性
- 推理增强:提供示例以鼓励模型遵循清晰的推理流程,学习将问题分解为逻辑步骤。
- 任务覆盖:包含基础、中级、高级和复杂推理任务。
- 语言专注:所有示例均使用自然的印尼语编写,确保与现实世界用法和语言模式保持一致。
主要用途
该数据集针对以下场景进行了优化:
- 训练 / 预训练
- 监督式微调 (SFT)
- 指令微调
- 推理增强流程
数据集结构
数据集被划分为多个子集,目前仍在持续扩展中。
| 配置名称 | 默认 | 数据文件 |
|---|---|---|
2026-05-15 |
是 | data/id-ID/2026-05-15.parquet |
2026-05-16 |
否 | data/id-ID/2026-05-16.parquet |
2026-05-17 |
否 | data/id-ID/2026-05-17.parquet |
2026-05-18_MathArena_aime_2026 |
否 | data/id-ID/2026-05-18_MathArena_aime_2026.parquet |
2026-05-18 |
否 | data/id-ID/2026-05-18.parquet |
2026-05-19 |
否 | data/id-ID/2026-05-19.parquet |
- 样本数量:少于 1,000 条 (n<1K)
- 任务类别:文本生成,问答
- 许可证:GPL-2.0
数据字段
数据集中每条记录包含以下必需的字段:
- user: 用户的输入问题或指令。
- reasoning: 模型思考或推理过程的文本。
- assistant: 模型基于推理过程生成的最终回答。
引用
bibtex @misc{openbahasa-cot, title={OpenBahasa-CoT}, author={hadadxyz}, year={2026}, url={https://huggingface.co/datasets/hadadxyz/OpenBahasa-CoT} }




