Bangla Idiom Misuse Detection Dataset: 15,670 Annotated Sentences with Usage Labels and Reasoning Explanations
收藏资源简介:
The dataset constructed for this study represents the first manually annotated corpus designed specifically for Bangla idiom usage validation and explainable reasoning generation. The corpus comprises 15,670 sentences spanning 1,316 distinct Bangla idioms (বাগধারা), with an average of 11.91 sentences per idiom to ensure comprehensive contextual coverage. Each instance was labeled for binary usage correctness (correct/incorrect) and augmented with a human-written reasoning explanation articulating the linguistic justification for the assigned label. The corpus exhibits a near-perfectly balanced label distribution, with 7,836 instances (50.01%) representing correct usage and 7,834 instances (49.99%) representing incorrect usage, preventing class imbalance bias during training and evaluation. The 1,316 idioms were collected from the publicly available Bangla Wikipedia page "বাংলা বাগধারার তালিকা", providing a standardized and widely referenced inventory. All idioms listed at the time of collection were included to avoid selective bias, with semantic diversity naturally emerging from the compilation and spanning emotional states, interpersonal relations, cognitive conditions, social behavior and everyday experiences.



