Aynkader/Balochi-dataset
收藏资源简介:
Balochi Latin Syáhag平行语料库与语法数据集是一个开源语言数据集,旨在促进Balochi Latin脚本(Balochi Latin Syáhag)集成到机器翻译平台和大型语言模型(LLMs)中。它包含一个平行翻译语料库,有18,000个已验证的英语与Balochi拉丁脚本之间的句对,以及一份官方正字法参考手册,详细说明了标准语法、变音规则和结构指南。数据集遵循Balochi拉丁正字法系统的核心语言规则,如排除V、F、Q和X字符,并系统使用元音变音符号(如á、é)以确保语音一致性。
Balochi Latin Syáhag Parallel Corpus & Grammar Dataset is an open-source linguistic dataset designed to facilitate the integration of the Balochi Latin script (Balochi Latin Syáhag) into machine translation platforms and large language models (LLMs). It includes a parallel translation corpus with 18,000 verified sentence pairs between English and Balochi Latin Script, and an official orthographic reference manual detailing standard syntax, diacritic rules, and structural guidelines. The dataset adheres to core linguistic rules of the Balochi Latin Syáhag system, such as excluding the characters V, F, Q, and X, and systematically implementing vowel diacritics (e.g., á, é) for phonological consistency.




