BengVoice: A Stratified Dataset of Code-Mixed Bengali-English Voice Commands for Intent Classification in Conversational AI Systems
收藏资源简介:
This dataset presents a meticulously curated benchmark collection of 1,200 Bengali voice assistant utterances for intent classification research in conversational AI systems. BengVoice addresses the critical gap in Natural Language Understanding resources for Bengali, one of the world's most widely spoken languages with over 230 million speakers, yet significantly underrepresented in publicly available language technology datasets. The dataset comprises utterances across 10 fundamental voice assistant intent categories: weather queries, time queries, alarm setting, news requests, music playback, phone calls, messaging, translation, calculations, and general knowledge questions. Each intent category contains exactly 120 samples, ensuring perfect class balance. All 1,200 utterances are unique with zero duplicates. A distinguishing feature is authentic code-mixing behaviour—natural integration of English words within Bengali speech. Analysis reveals 290 samples (24.2%) contain code-mixed content, with patterns reflecting genuine usage: technical domains like alarm setting show 71.7% code-mixing, while traditional domains show minimal mixing (0.8%). This reflects natural speech patterns of urban Bengali speakers in Bangladesh. The dataset incorporates cultural authenticity through references to Bangladeshi locations (Dhaka, Chittagong, Sylhet), local media (Prothom Alo, Kaler Kantho), and cultural elements specific to Bangladesh, ensuring real-world usage scenarios for Bengali-speaking populations. For robust evaluation, the dataset provides stratified 5-fold cross-validation splits. Each fold contains exactly 240 samples with 24 per intent, maintaining perfect balance. This stratification enables fair model comparison and supports multiple evaluation methodologies including traditional machine learning, deep learning, retrieval-augmented generation (RAG), and few-shot prompting. Baseline validation experiments using TF-IDF vectorization with character-level n-grams and Logistic Regression achieved mean accuracy of 93.92% (±0.50%) across 5-fold cross-validation, with fold accuracies from 93.33% to 94.58%. Per-intent performance ranged from 81.67% (news requests) to 100% (translation), establishing clear benchmarks and validating dataset quality. The dataset is provided in multiple formats: complete datasets in JSON and CSV (with and without fold labels), individual fold files for pre-separated evaluation. No proprietary software required. This resource enables Bengali voice assistant development, intent classification benchmarking, code-mixing investigation, cross-lingual transfer learning, multilingual NLU systems, and low-resource language processing. Released under Creative Commons Attribution 4.0 International (CC BY 4.0) license for maximum research impact.
本数据集精心构建了一套基准合集,包含1200条孟加拉语语音助手话语,用于对话式人工智能系统中的意图分类研究。孟加拉语是全球使用人数超2.3亿的主流语言之一,但在公开可用的语言技术数据集中却占比极低,BengVoice数据集填补了孟加拉语自然语言理解(Natural Language Understanding, NLU)资源的关键空白。 该数据集涵盖10类基础语音助手意图分类任务:天气查询、时间查询、闹钟设置、新闻请求、音乐播放、拨打电话、发送消息、翻译服务、数值计算以及常识问答。每类意图恰好包含120条样本,确保类别分布完全均衡。全部1200条话语均为唯一样本,无任何重复。 该数据集的显著特征是包含真实的语码混合现象——英语词汇自然融入孟加拉语口语之中。分析显示,其中290条样本(占比24.2%)包含语码混合内容,其模式符合真实使用场景:闹钟设置等技术类场景的语码混合占比达71.7%,而传统类场景的混合占比仅为0.8%。这一特征精准反映了孟加拉国城市孟加拉语使用者的自然口语习惯。 数据集通过融入孟加拉国本土元素保障文化真实性,其中包含孟加拉国本地地名(达卡、吉大港、锡尔赫特)、本土媒体(Prothom Alo、Kaler Kantho)以及孟加拉国特有的文化元素,贴合孟加拉语使用者的真实使用场景。 为支持稳健的模型评估,数据集提供分层5折交叉验证划分方案。每一折均包含240条样本,每类意图对应24条样本,始终保持类别均衡。这种分层划分方法可实现公平的模型对比,并支持多种评估方法论,包括传统机器学习、深度学习、检索增强生成(Retrieval-Augmented Generation, RAG)以及少样本提示。 采用基于字符级n-gram的TF-IDF向量化结合逻辑回归(Logistic Regression)的基线验证实验,在5折交叉验证中取得了93.92%(±0.50%)的平均准确率,各折准确率介于93.33%至94.58%之间。单意图准确率范围为81.67%(新闻请求)至100%(翻译服务),这一结果确立了明确的基准指标,并验证了数据集的质量。 本数据集提供多种存储格式:包含完整数据集的JSON与CSV文件(带/不带折次标签),以及用于预拆分评估的独立折次文件。无需专有软件即可使用。 该资源可用于孟加拉语语音助手开发、意图分类基准测试、语码混合研究、跨语言迁移学习、多语言自然语言理解系统以及低资源语言处理研究。本数据集采用知识共享署名4.0国际(Creative Commons Attribution 4.0 International, CC BY 4.0)许可协议发布,以最大化其研究影响力。




