Natural Devanagari Character Digit Audio Dataset - DEVA-AUDIO
收藏资源简介:
This dataset presents a collection of human-recorded audio samples of Devanagari characters designed for speech processing and CAPTCHA-based security applications. The dataset consists of isolated character-level recordings, including vowels, consonants, and digits, spoken by multiple contributors under varied recording conditions. Unlike synthetic speech datasets, all recordings are natural, capturing real-world variations in pronunciation, pitch, and background noise. The dataset aims to support research in automatic speech recognition (ASR), human-computer interaction, and accessible CAPTCHA systems, particularly for Indian language users. With over 10,000 audio samples, this dataset provides a valuable resource for developing robust and inclusive speech-based systems.A Python-based processing pipeline is implemented for audio format conversion (M4A to MP3), pause-based segmentation, and systematic organization of audio samples. A total of 10,675 validated audio samples are generated, including 7,108 character samples and 3,567 digit samples.
本数据集收录了人类录制的天城文(Devanagari)字符音频样本,专为语音处理与基于验证码(CAPTCHA)的安全应用场景设计。该数据集包含独立的字符级录音样本,涵盖元音、辅音与数字,由多名贡献者在多样的录音环境下完成录制。与合成语音数据集不同,所有录音均为自然语音,完整捕捉了发音、音调与背景噪声的真实场景差异。本数据集旨在助力自动语音识别(Automatic Speech Recognition,ASR)、人机交互及无障碍验证码系统相关研究,尤其面向印度语言使用者。本数据集拥有超10000条音频样本,可为构建鲁棒性强且具备包容性的语音交互系统提供宝贵资源。项目搭建了基于Python的处理流水线,可实现音频格式转换(M4A转MP3)、基于停顿的分段处理,以及音频样本的系统化整理。最终生成共计10675条经过验证的音频样本,其中包含7108条字符样本与3567条数字样本。



