遇见数据集

Sandroeth/code-bilingual-it

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

Code Bilingual IT是一个基于Alpaca格式的合成数据集,专为训练语言模型以遵循指令并生成简单代码而设计。该数据集包含每个条目的印尼语和英语双语版本,覆盖8种编程语言(Python、JavaScript、C++、Bash、SQL、JSON、YAML、Markdown),总计39,000个条目。每个条目包括指令、输入上下文、输出(含简短解释和代码块)、编程语言、难度级别(初级、中级、高级)和语言区域(英语或印尼语)。数据集规模在10K到100K之间,适用于文本生成和语言建模任务。

Code Bilingual IT is a synthetic dataset based on the Alpaca format, designed to train language models to follow instructions and generate simple code. Each entry is available in two languages, Indonesian and English, covering 8 programming languages with a total of 39,000 entries. The dataset includes fields such as instruction, input, output (with brief explanation and code block), language, difficulty (beginner, intermediate, advanced), and locale (en or id). It is suitable for text-generation and language-modeling tasks, with a size category of 10K to 100K.

提供机构:
Sandroeth
二维码
社区交流群
二维码
科研交流群
商业服务