VisualVoiceTherapy (VVT) Dataset v1: Exemplary Audio-Visual Voice Therapy Data
收藏官方服务:
资源简介:
The VisualVoiceTherapy (VVT) dataset is an audiovisual voice therapy dataset comprising 818 valid videos (14.53 hours, 40.9 GB) from 53 participants performing structured voice therapy exercises. From these recordings, video frames were extracted and annotated to create a domain-specific keypoint dataset. The dataset includes 39 facial keypoints and 21 upper-body keypoints. In total, 8,287 frames were annotated for facial landmarks and 2,097 frames for upper-body landmarks. Data collection took place in Germany; therefore, all spoken content is provided in German.
提供机构:
Zenodo创建时间:
2026-01-19



