官方服务:
资源简介:
Bangla Lip2Text 5-Sec Video Dataset
应用场景:
相关数据集
Bengali Voice Commerce Audio Dataset for Automatic Speech Recognition
This meticulously curated dataset presents a comprehensive collection of Bengali speech recordings meticulously compiled for the advancement of voice-enabled commerce applications. Consisting of 3344
Mendeley Data2024-03-27 更新110
MODALITY corpus - SPEAKER 35 - SEQUENCE S5
The MODALITY corpus is one of the multimodal database of word recordings in English. It consists of over 30 hours of multimodal recordings. The database contains high-resolution, high-framerate stereo
DataCite Commons2026-04-10 更新60
BanglaMood: A Rich Audio Dataset for Eight Emotion Recognition in Bangla Speech
The BanglaMOOD dataset comprises 3,220 audio files in WAV format, each capturing vocal expressions of human emotion in Bangla. The dataset features eight emotion categories: Angry, Happy, Sad, Neutral
Mendeley Data30
Bayanno
This dataset is published for Bengali continuous speech recognition. The dataset has three files "Script Files" contain the Bengali text for continuous speech and "Speech Files" contain the recorded a
Mendeley Data2021-02-22 更新70
rv-4d54f178b0
本数据集为俄语唇读语料库,专门用于视觉语音识别(VSR)任务,仅包含视频片段(无音频)。数据规模为229,494个视频片段,总时长约690小时(截至2026年9月3日),是目前公开最大的俄语VSR数据集(对比MuAViC仅49小时)。数据来源分为两部分:一部分(160,921个片段,306小时)来自牛津VGG的MultiVSR开放列表,经本数据集提供的流水线切分而成;另一部分(68,573个片段,
Hugging Face2026-09-03 更新40



