LRW-Persian
收藏资源简介:
LRW-Persian是迄今为止最大的波斯语单词级唇读数据集,包含743个目标单词和超过414,000个视频样本,这些样本是从超过67个电视节目的1,900多个小时的镜头中提取的。该数据集旨在作为一个基准资源,提供speaker-disjoint训练和测试分割,广泛的地域和方言覆盖,以及丰富的每个剪辑元数据,包括头部姿态、年龄和性别。为了确保大规模数据质量,我们建立了一个全自动的端到端策展流程,包括基于自动语音识别(ASR)的转录、主动说话者定位、质量过滤和姿态/面具筛选。
LRW-Persian is the largest Persian word-level lip-reading dataset to date. It contains 743 target words and over 414,000 video samples, which are extracted from more than 1,900 hours of footage sourced from over 67 television programs. This dataset is intended to serve as a benchmark resource, providing speaker-disjoint training and test splits, extensive geographic and dialectal coverage, and comprehensive per-clip metadata including head pose, age, and gender. To guarantee the quality of the large-scale dataset, we developed a fully automated end-to-end curation pipeline that encompasses ASR-based transcription, active speaker localization, quality filtering, and pose/mask screening.



