AiRukua/Geser_Indo_En
收藏资源简介:
该数据集是一个手工创建的三语平行语料库,包含约7,000个句子对,涵盖Geser语(巴哈萨塞拉姆的Geser方言)、印尼语(标准印尼语)和英语(标准英语)。数据由Geser语母语者完全手工编写、审核和验证,未使用任何机器翻译或自动对齐技术。该语料库旨在记录和保护印度尼西亚马鲁古群岛塞兰岛地区的一种濒危语言——Geser语,这是一种使用人数较少且面临代际传承危机的语言。数据集是公开可用的首批Geser语结构化语言资源之一,适用于低资源机器翻译、语言文档记录、语言学分析和NLP基准测试等用途。
This dataset is a manually curated trilingual parallel corpus containing approximately 7,000 sentence pairs across Geser (a dialect of Bahasa Seram), Indonesian (standard Bahasa Indonesia), and English (standard English). The data was entirely handcrafted by native speakers of Geser, with each sentence written, reviewed, and validated by fluent speakers—no machine translation or automated alignment was used. The corpus aims to document and preserve Geser, an endangered language spoken in the Seram island region of Maluku, Indonesia, which has a small, aging speaker population and declining intergenerational transmission. It is one of the first structured linguistic resources for Geser made publicly available, intended for low-resource machine translation, language documentation, linguistic analysis, and NLP benchmarking.



