遇见数据集

Bangla Idiom Misuse Detection Dataset: 15,670 Annotated Sentences with Usage Labels and Reasoning Explanations

收藏
Zenodo2026-04-03 更新2026-05-26 收录
官方服务:

资源简介:

The dataset constructed for this study represents the first manually annotated corpus designed specifically for Bangla idiom usage validation and explainable reasoning generation. The corpus comprises 15,670 sentences spanning 1,316 distinct Bangla idioms (বাগধারা), with an average of 11.91 sentences per idiom to ensure comprehensive contextual coverage. Each instance was labeled for binary usage correctness (correct/incorrect) and augmented with a human-written reasoning explanation articulating the linguistic justification for the assigned label. The corpus exhibits a near-perfectly balanced label distribution, with 7,836 instances (50.01%) representing correct usage and 7,834 instances (49.99%) representing incorrect usage, preventing class imbalance bias during training and evaluation. The 1,316 idioms were collected from the publicly available Bangla Wikipedia page "বাংলা বাগধারার তালিকা", providing a standardized and widely referenced inventory. All idioms listed at the time of collection were included to avoid selective bias, with semantic diversity naturally emerging from the compilation and spanning emotional states, interpersonal relations, cognitive conditions, social behavior and everyday experiences.

提供机构:
Zenodo
创建时间:
2026-04-03
二维码
社区交流群
二维码
科研交流群
商业服务