遇见数据集

Layers and parameters of the 2DS-CNN model.

收藏
Figshare2024-04-26 更新2026-04-28 收录
官方服务:

资源简介:

Digital speech recognition is a challenging problem that requires the ability to learn complex signal characteristics such as frequency, pitch, intensity, timbre, and melody, which traditional methods often face issues in recognizing. This article introduces three solutions based on convolutional neural networks (CNN) to solve the problem: 1D-CNN is designed to learn directly from digital data; 2DS-CNN and 2DM-CNN have a more complex architecture, transferring raw waveform into transformed images using Fourier transform to learn essential features. Experimental results on four large data sets, containing 30,000 samples for each, show that the three proposed models achieve superior performance compared to well-known models such as GoogLeNet and AlexNet, with the best accuracy of 95.87%, 99.65%, and 99.76%, respectively. With 5-10% higher performance than other models, the proposed solution has demonstrated the ability to effectively learn features, improve recognition accuracy and speed, and open up the potential for broad applications in virtual assistants, medical recording, and voice commands.

数字语音识别是一项极具挑战性的研究课题,其要求模型能够学习语音信号的复杂特征,涵盖频率、基音、强度、音色与旋律等维度,而传统识别方法在这类任务中往往存在识别精度不足的问题。本文提出三种基于卷积神经网络(Convolutional Neural Network)的解决方案以攻克该难题:一维卷积神经网络(1D-CNN)被设计为直接从原始数字语音数据中学习特征;二维频谱卷积神经网络(2DS-CNN)与二维梅尔频谱卷积神经网络(2DM-CNN)架构更为复杂,它们通过傅里叶变换将原始语音波形转换为频谱图像,进而提取关键特征。在四个各包含30000个样本的大型数据集上开展的实验结果表明,所提三种模型的性能均优于GoogLeNet、AlexNet等经典深度学习模型,三者的最优识别准确率分别为95.87%、99.65%与99.76%。相较于其他同类模型,所提方案的性能提升5%至10%,其不仅展现出高效学习语音特征、提升识别准确率与推理速度的能力,更为其在虚拟助手、医疗录音、语音指令等领域的广泛应用开辟了广阔前景。

创建时间:
2024-04-26
二维码
社区交流群
二维码
科研交流群
商业服务