Wilder Active Speaker Detection (WASD) Dataset
收藏资源简介:
WASD数据集是由葡萄牙贝拉内大学电信研究所创建,旨在解决当前主动说话人检测模型在非受控环境下的局限性。该数据集包含164个视频,分为5个难度递增的类别,从最佳条件到监控设置,专门设计来测试音频和面部特征的挑战。WASD数据集不仅包含面部和音频数据,还提供了身体数据注释,以促进开发利用身体信息补充面部和音频数据的模型。此数据集的应用领域包括评估现有模型在野外条件下的性能,以及推动主动说话人检测技术的进步。
The WASD Dataset was created by the Instituto de Telecomunicações, University of Beira Interior, Portugal, aiming to address the limitations of current active speaker detection models in uncontrolled environments. This dataset includes 164 videos categorized into five difficulty-increasing classes, spanning from optimal conditions to surveillance settings, and is specifically designed to test the challenges posed by audio and facial characteristics. Beyond facial and audio data, the WASD Dataset also offers body data annotations, enabling the development of models that leverage body information to complement facial and audio modalities. The applications of this dataset cover evaluating the performance of existing models in wild environments, as well as advancing active speaker detection technologies.




