bengaliAI/bethik-lexicon
收藏资源简介:
bethik-lexicon是一个用于孟加拉语语法错误检测(GED)的数据集,包含孟加拉语单词列表和字素到拉丁字母的映射。它用于Bhashabhrom/bethik GED管道,支持拼写检查和语法错误检测任务。数据集包括两个主要文件:words.csv(源词典,约667 MiB,包含单词列)和gmap.json(用于资源创建脚本的映射文件)。语言为孟加拉语(bn),涉及标签包括bengali、bangla、spell-check和grammatical-error-detection。
The bethik-lexicon is a dataset dedicated to Bengali grammatical error detection (GED). It contains a list of Bengali words and a grapheme-to-Latin letter mapping. This dataset is employed in the Bhashabhrom/bethik GED pipeline to support both spell-check and grammatical error detection tasks. The dataset comprises two core files: words.csv (a source lexicon with a size of approximately 667 MiB, including a word column) and gmap.json (a mapping file for resource creation scripts). The language of the dataset is Bengali (bn), and its associated tags include bengali, bangla, spell-check, and grammatical-error-detection.




