BD-GRF6: A Six-Class Bangladeshi Speech Dataset for Real and Fake Classification Across Male, Female, and Third Gender
收藏资源简介:
BD-GRF6 is a structured six-class Bangladeshi speech dataset developed to support research in gender-aware real and fake voice classification. The dataset includes speech recordings from three gender categories: Male, Female, and Third Gender. Each gender category contains both authentic (real) and manipulated (fake/spoofed) speech samples, resulting in six distinct classes: Male Real, Male Fake, Female Real, Female Fake, Third Gender Real, and Third Gender Fake. The dataset comprises a total of 13,566 audio files with an overall duration of approximately 18.7 hours. The recordings are organized into clearly separated class-wise directories to facilitate supervised machine learning and deep learning experiments. The dataset is designed to support tasks such as real vs. fake voice detection, gender-aware spoof detection, and multi-class speech classification. All speech samples are in Bangla (Bengali) and represent Bangladeshi speakers. The dataset aims to promote research in speech forensics, anti-spoofing systems, biometric security, and inclusive gender-based voice analysis. By including third gender speech data, BD-GRF6 contributes toward more inclusive and representative speech datasets. This dataset can be used for academic research, benchmarking classification models, and developing robust real and fake speech detection systems.



