遇见数据集

LipBengal Dataset

收藏
DataCite Commons2025-06-01 更新2024-08-19 收录
官方服务:

资源简介:

We introduce the "LipBengal" dataset, marking a significant advancement in the field of Bengali lip-reading and visual speech recognition research. This dataset addresses a critical gap in the research landscape. Despite Bengali being the seventh most spoken language globally, with over 265 million speakers, it has been largely underrepresented in this domain.LipBengal offers a comprehensive resource for researchers, comprising visual data from 150 speakers across 73 classes, covering Bengali phonemes, alphabets, and symbols. Captured under diverse and uncontrolled conditions, LipBengal is the most extensive Bengali lip-reading dataset to date. Detailed annotations, ranging from phoneme-level classifications to full sentence constructions, further enhance its value. The dataset's thorough coverage of Bengali phonemes captures the nuances of lip movements associated with distinct sounds.This rich resource holds promise for training accurate lip-reading models, with applications in accessibility improvements, enhanced speech recognition, silent speech interfaces, and linguistic research. The diversity in speaker backgrounds ensures broader representability of Bengali pronunciation patterns, while meticulous annotation and curation processes guarantee quality and reliability. LipBengal is a valuable asset for researchers and developers working in Bengali lip-reading and visual speech recognition.<br>The LipBengal dataset can be accessed through:<br>https://drive.google.com/drive/folders/1CgOg35Cfs3H6-vHmG11LDt0qlS6q--jt?usp=sharing

我们提出了LipBengal数据集,为孟加拉语唇读与视觉语音识别研究领域带来了重要进展。该数据集填补了该领域的关键研究空白。尽管孟加拉语是全球第七大通用语言,使用者超2.65亿,但该语言在唇读研究领域的代表性仍严重不足。<br>LipBengal为研究人员提供了全面的研究资源,涵盖来自150位发音者的73个类别的视觉数据,内容覆盖孟加拉语音素、字母与各类符号。该数据集采集自多样化且非受控的环境,是目前规模最大的孟加拉语唇读数据集。其附带从音素层级分类到完整语句构建的精细化标注,进一步提升了数据集的研究价值。<br>该数据集对孟加拉语音素的全面覆盖,能够精准捕捉不同发音对应的唇动细节。这一丰富的资源可用于训练高精度唇读模型,其应用场景涵盖无障碍体验优化、增强型语音识别、静默语音交互界面以及语言学研究等领域。发音者背景的多样性,确保了孟加拉语发音模式的广泛代表性;而精细化的标注与整理流程,则保障了数据集的质量与可靠性。对于从事孟加拉语唇读与视觉语音识别研究的人员与开发者而言,LipBengal数据集是极具价值的研究资产。<br>LipBengal数据集可通过以下链接获取:<br>https://drive.google.com/drive/folders/1CgOg35Cfs3H6-vHmG11LDt0qlS6q--jt?usp=sharing

提供机构:
figshare
创建时间:
2024-06-11
搜集汇总
数据集介绍
LipBengal Dataset 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务