遇见数据集

Bangla Idiom Misuse Detection Dataset: 15,670 Annotated Sentences with Usage Labels and Reasoning Explanations

收藏
Zenodo2026-04-03 更新2026-05-26 收录
官方服务:

资源简介:

The dataset constructed for this study represents the first manually annotated corpus designed specifically for Bangla idiom usage validation and explainable reasoning generation. The corpus comprises 15,670 sentences spanning 1,316 distinct Bangla idioms (বাগধারা), with an average of 11.91 sentences per idiom to ensure comprehensive contextual coverage. Each instance was labeled for binary usage correctness (correct/incorrect) and augmented with a human-written reasoning explanation articulating the linguistic justification for the assigned label. The corpus exhibits a near-perfectly balanced label distribution, with 7,836 instances (50.01%) representing correct usage and 7,834 instances (49.99%) representing incorrect usage, preventing class imbalance bias during training and evaluation. The 1,316 idioms were collected from the publicly available Bangla Wikipedia page "বাংলা বাগধারার তালিকা", providing a standardized and widely referenced inventory. All idioms listed at the time of collection were included to avoid selective bias, with semantic diversity naturally emerging from the compilation and spanning emotional states, interpersonal relations, cognitive conditions, social behavior and everyday experiences.

本研究构建的数据集是首个专门面向孟加拉语成语(Bangla idiom)使用验证与可解释推理生成的人工标注语料库。该语料库包含15670个句子,涵盖1316种不同的孟加拉语成语(বাগধারা),每个成语平均对应11.91个句子,以确保实现全面的上下文覆盖。每个样本均被标注了二元使用正确性标签(正确/错误),并附带人工撰写的推理解释,阐明该标签对应的语言学依据。该语料库的标签分布近乎完美平衡,其中7836个样本(占比50.01%)为正确使用场景,7834个样本(占比49.99%)为错误使用场景,可有效避免训练与评估过程中出现类别不平衡偏差。本次采集的1316个成语均来自公开可用的孟加拉语维基百科页面“বাংলা বাগধারার তালিকা”,该来源为标准化且被广泛引用的孟加拉语成语清单。采集时收录了该页面当时收录的全部成语,以规避选择性偏差,最终生成的语料库自然具备丰富语义多样性,覆盖情绪状态、人际关系、认知状态、社会行为及日常体验等多个范畴。

提供机构:
Zenodo
创建时间:
2026-04-03
二维码
社区交流群
二维码
科研交流群
商业服务