Code-170k-nuer
收藏资源简介:
Code-170k-nuer 是一个包含 176,999 个编程对话的数据集,这些对话被翻译成了努尔语,旨在让努尔语使用者能够接触并学习编程。数据集中的对话涵盖了各种编程概念,适用于训练努尔语编程助手、构建教育工具、研究多语言代码生成等多种用途。数据集的结构为一个包含对话轮次的列表,每个轮次包括 'from'(发言者)和 'value'(消息内容)。
Code-170k-nuer is a dataset consisting of 176,999 programming conversations translated into the Nuer language, designed to provide Nuer speakers with access to and facilitate their learning of programming. The conversations within this dataset cover a broad spectrum of programming concepts, and support multiple use cases including training Nuer-language programming assistants, developing educational tools, and researching multilingual code generation. The dataset is structured as a list of conversation turns, with each turn containing two fields: 'from' (indicating the speaker) and 'value' (representing the message content).
Code-170k-nuer 数据集概述
基本信息
- 数据集名称: Code-170k-nuer
- 发布年份: 2025
- 发布平台: Hugging Face
- 许可证: Apache 2.0
- 语言: 努尔语(Nuer)
数据集规模
- 训练集样本数量: 176,999
- 训练集大小: 383,754,208字节
- 下载大小: 191,877,104字节
- 规模分类: 100K<n<1M
数据特征
数据结构
- 主要字段: conversations
- 对话轮次结构:
- from: 说话者标识("human"或"gpt")
- value: 努尔语消息内容
数据格式示例
python { "conversations": [ { "from": "human", "value": "[努尔语问题]" }, { "from": "gpt", "value": "[努尔语回答]" } ] }
数据集特点
- 数据来源: 基于glaiveai/glaive-code-assistant-v2翻译
- 语言特性: 纯努尔语编程对话
- 对话类型: 多轮对话
- 内容主题: 编程概念、算法、数据结构、调试、最佳实践等
任务类别
- 文本生成
- 问答系统
应用场景
- 努尔语编程助手训练
- 努尔开发者教育工具开发
- 多语言代码生成研究
- 努尔语编程教程创建
- 低资源语言AI开发支持
技术标签
- code
- programming
- nus
- nuer
- african-languages
- low-resource
- multilingual
- instruction-tuning
引用格式
bibtex @dataset{code170k_nuer, title={Code-170k-nuer: Programming Conversations in Nuer}, year={2025}, publisher={Hugging Face}, url={https://huggingface.co/datasets/michsethowusu/Code-170k-nuer} }




