Balochi Latin Syáhag Parallel Corpus & Grammar Dataset
收藏资源简介:
该仓库包含一个高质量、经过验证的语言学数据集,旨在促进Balochi Latin脚本(Balochi Latin Syáhag)集成到机器翻译平台、本地化神经网络和开源计算模型中。数据集内容包括一个正式的拼写参考手册(PDF),详细描述了标准语法、变音符号规则和形态结构,以及一个包含18,000个句子的平行语料库(CSV),其中包含双语句子对,用于机器学习模型对齐。数据集遵循Balochi Latin Syáhag系统的核心语言约束,如排除字符V、F、Q和X,并系统实施标准化元音重音(á, é, ó)以确保跨不同Balochi方言的音系一致性。该数据集旨在轻松集成到开源基准测试框架、本地化管道和平台翻译训练模型(如Google Translate请求)中,干净的CSV翻译对允许直接计算标记化,预处理最少。数据集以开放许可提供,支持数字语言包容性。
This repository contains a high-quality, validated linguistic dataset developed to promote the integration of the Balochi Latin script (Balochi Latin Syáhag) into machine translation platforms, localized neural networks, and open-source computational models. The dataset includes a formal spelling reference manual (PDF) that details standard grammar, diacritic rules, and morphological structure, as well as a parallel corpus (CSV) consisting of 18,000 bilingual sentence pairs for machine learning model alignment. The dataset adheres to the core linguistic constraints of the Balochi Latin Syáhag system, such as excluding the characters V, F, Q, and X, and systematically implements standardized vowel diacritics (á, é, ó) to ensure phonological consistency across different Balochi dialects. This dataset is designed for easy integration into open-source benchmarking frameworks, localization pipelines, and translation training models for platforms such as Google Translate; the clean CSV sentence pairs allow direct tokenization computation with minimal preprocessing. The dataset is released under an open license to support digital linguistic inclusion.
数据集概述
该数据集是一个高质量、经过验证的语言学数据集,旨在支持俾路支语拉丁字母(Balochi Latin Syáhag)在机器翻译平台、本地化神经网络及开源计算模型中的集成应用。
数据集内容
- PDF参考手册:文件名为
Balóchiay Látini Syáhagay Rahband.pdf,是一份正式的正字法参考手册,详细说明了标准语法、变音符号规则及形态结构。 - 双语平行语料:文件名为
dataset.csv,包含 18,000 句 双语平行句对,按句对排列,便于机器学习模型对齐。
正字法规则与约束
训练分词器或语言模型时,必须遵守以下核心语言学约束:
- 禁止字符:字母 V、F、Q、X 在该正字法中严格禁用,不出现于任何真实词汇中。
- 变音符号与元音:系统性地使用标准化元音变音符号(如 á、é、ó),以确保跨方言的语音一致性。
计算应用目的
该数据结构化设计,可轻松集成到开源基准测试框架、本地化流程及平台翻译训练模型(如 Google 翻译请求)中,清洗后的 CSV 翻译对支持直接计算分词,预处理需求极低。
许可与联系
该数据集以开放形式提供,旨在促进数字语言包容性。有关查询、协作扩展或进一步的语言验证,可通过仓库的 Issue 功能或联系维护者进行沟通。




