Bangla Code Instruction dataset
收藏资源简介:
本数据集为Bangla语言编程领域提供了首个全面的代码指令数据集,包含30万个指令-代码对,用于编程领域的适应。数据集由三部分组成:Bangla-Code-Instruct-SI、Bangla-Code-Instruct-Syn和Bangla-Code-Instruct-TE,分别包含10万个指令-代码对,涵盖自我指导、合成生成和机器翻译三种方法。数据集旨在促进低资源语言代码生成模型的发展,并通过开源资源推动Bangla语言代码生成领域的研究。
This dataset is the first comprehensive code instruction dataset for the Bangla programming domain, containing 300,000 instruction-code pairs for programming domain adaptation. The dataset consists of three subsets: Bangla-Code-Instruct-SI, Bangla-Code-Instruct-Syn, and Bangla-Code-Instruct-TE, each with 100,000 instruction-code pairs, covering three data construction methods: self-instruction, synthetic generation, and machine translation. This dataset aims to facilitate the development of code generation models for low-resource languages, and promote research in the Bangla code generation domain via open-source resources.




