遇见数据集

Wilder Active Speaker Detection (WASD) Dataset

收藏
arXiv2023-03-09 更新2024-06-21 收录
数据链接:
官方服务:

资源简介:

WASD数据集是由葡萄牙贝拉内大学电信研究所创建,旨在解决当前主动说话人检测模型在非受控环境下的局限性。该数据集包含164个视频,分为5个难度递增的类别,从最佳条件到监控设置,专门设计来测试音频和面部特征的挑战。WASD数据集不仅包含面部和音频数据,还提供了身体数据注释,以促进开发利用身体信息补充面部和音频数据的模型。此数据集的应用领域包括评估现有模型在野外条件下的性能,以及推动主动说话人检测技术的进步。

The WASD Dataset was created by the Instituto de Telecomunicações, University of Beira Interior, Portugal, aiming to address the limitations of current active speaker detection models in uncontrolled environments. This dataset includes 164 videos categorized into five difficulty-increasing classes, spanning from optimal conditions to surveillance settings, and is specifically designed to test the challenges posed by audio and facial characteristics. Beyond facial and audio data, the WASD Dataset also offers body data annotations, enabling the development of models that leverage body information to complement facial and audio modalities. The applications of this dataset cover evaluating the performance of existing models in wild environments, as well as advancing active speaker detection technologies.

提供机构:
电信研究所
创建时间:
2023-03-09
搜集汇总
数据集介绍
Wilder Active Speaker Detection (WASD) Dataset 数据集图片
背景与挑战
背景概述
Wilder Active Speaker Detection (WASD) 数据集是一个用于主动说话人检测的基准数据集,通过设计5个从最优条件到监控设置的类别,逐步增加音频和面部数据的挑战性,以评估模型在真实复杂场景下的性能。该数据集包含预处理后的音频、视频和标注文件,总大小约300GB,旨在推动ASD技术在实际应用中的发展,并已在相关学术论文中得到验证。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务