BiST
收藏资源简介:
BiST是由库尔纳工程技术大学与肯尼索州立大学联合构建的首个孟加拉语-英语双语语法标注语料库,包含30,534条句子(英语17,465条,孟加拉语13,069条),数据来源于维基百科和日常对话文本。该数据集通过多阶段预处理和三位独立标注者的维度级Fleiss’ Kappa一致性验证(结构标注κ=0.82,时态标注κ=0.88),标注了句法结构(简单/复合/复杂/复杂复合句)和时态(现在/过去/将来)双维度信息。其核心应用涵盖语法建模、跨语言表示学习及教育NLP系统开发,旨在解决低资源语言场景下双语语法标注数据匮乏的问题。
BiST is the first Bengali-English bilingual syntactically annotated corpus jointly constructed by Khulna University of Engineering & Technology and Kennesaw State University. It contains a total of 30,534 sentences, with 17,465 in English and 13,069 in Bengali, sourced from Wikipedia and daily conversational texts. This dataset has undergone multi-stage preprocessing and was validated for inter-annotator consistency via dimension-wise Fleiss’ Kappa test by three independent annotators, with κ values of 0.82 for syntactic structure annotation and 0.88 for tense annotation. It is annotated with dual-dimensional information including syntactic structures (simple, compound, complex, and complex-compound sentences) and tense categories (present, past, and future). Its core applications cover grammatical modeling, cross-lingual representation learning, and educational NLP system development, aiming to address the shortage of bilingual syntactically annotated data in low-resource language scenarios.
BiST数据集概述
数据集基本信息
- 数据集名称:BiST
- 关联学术会议:LREC-2026
- 当前状态:相关论文已被LREC-2026会议录用
数据集关联文献
- 该数据集对应一篇学术论文,该论文已被LREC-2026会议收录,并计划发表于会议论文集中。




