lexeme-alignments
收藏资源简介:
lexeme-alignments数据集是一个多语言圣经词汇对齐资源,专注于从目标语言表面词形到源语言词位(lexeme)的映射关系。它覆盖12种语言(阿拉伯语、阿萨姆语、孟加拉语、英语、法语、豪萨语、印地语、印尼语、俄语、西班牙语、瑞典语、斯洛伐克语),通过三种对齐方法(eflomal统计对齐、gloss词典对齐、gapfill覆盖填充)挖掘得到。数据集采用词位锚定原则,以MACULA词位ID为核心标识,组织方式为可加性联合,保留所有对齐方法的原始证据,支持完整的来源追溯。它还支持多版本圣经译本合并,允许用户分析跨版本一致性。数据集质量经过Clear-Bible手动标注验证,在词位粒度上的top-1准确率约为89-92%,并包含四个配套参考资源文件。数据集本身采用CC0-1.0许可,但表面词形来自的源译本保留各自原始许可证。
The lexeme-alignments dataset is a multilingual biblical vocabulary alignment resource focusing on mapping from target language surface word forms to source language lexemes. It covers 12 languages (Arabic, Assamese, Bengali, English, French, Hausa, Hindi, Indonesian, Russian, Spanish, Swedish, Slovak) and is derived through three alignment methods: eflomal statistical alignment, gloss dictionary alignment, and gapfill coverage filling. The dataset employs a lexeme anchoring principle, using MACULA lexeme IDs as core identifiers, and is organized in an additive joint manner to retain original evidence from all alignment methods, ensuring full traceability. It supports merging multiple biblical translation versions, allowing users to analyze cross-version consistency as a strong confidence signal. Dataset quality is validated via Clear-Bible manual annotations, achieving a top-1 accuracy of approximately 89-92% at the lexeme granularity, and includes four supporting reference resource files. The dataset itself is licensed under CC0-1.0, but the surface word forms from source translations retain their original licenses.




