BeDeSpeech: A Bengali Dementia Speech Corpus
收藏资源简介:
This repository contains a Bengali Dementia Speech Dataset developed to support research on speech-based dementia detection and artificial intelligence for healthcare. The dataset consists of 740 speech recordings based on 5 basic needs (food, cloth, health, education, and shelter). Data were collected from native Bangladeshi Bengali speakers, including 370 recordings from individuals diagnosed with dementia and 370 recordings from cognitively healthy (non-dementia) participants. The recordings were collected under a standardized speech elicitation protocol to ensure consistency and facilitate comparative analysis between the two groups. The dataset is intended to support research in speech processing, machine learning, deep learning, natural language and speech technologies, digital health, and explainable artificial intelligence (XAI). It can be used to develop and evaluate automated systems for dementia detection, acoustic biomarker discovery, feature analysis, and interpretable AI models for clinical decision support. Researchers can utilize this dataset for a wide range of tasks, including audio preprocessing, acoustic feature extraction (e.g., MFCCs, pitch, jitter, shimmer, harmonic-to-noise ratio, and spectral features), feature selection, classification, deep learning, transfer learning, and explainable machine learning using methods such as SHAP and LIME. The dataset also provides opportunities to investigate speech characteristics associated with cognitive decline in a low-resource language setting. By making this dataset publicly available, we aim to encourage reproducible research, facilitate benchmarking of speech-based dementia detection algorithms, and promote the development of reliable and interpretable AI solutions for the early screening and assessment of dementia among Bengali-speaking populations.



