遇见数据集

VisualVoiceTherapy (VVT) Dataset v1: Exemplary Audio-Visual Voice Therapy Data

收藏
Zenodo2026-01-19 更新2026-05-26 收录
官方服务:

资源简介:

The VisualVoiceTherapy (VVT) dataset is an audiovisual voice therapy dataset comprising 818 valid videos (14.53 hours, 40.9 GB) from 53 participants performing structured voice therapy exercises. From these recordings, video frames were extracted and annotated to create a domain-specific keypoint dataset. The dataset includes 39 facial keypoints and 21 upper-body keypoints. In total, 8,287 frames were annotated for facial landmarks and 2,097 frames for upper-body landmarks. Data collection took place in Germany; therefore, all spoken content is provided in German.

提供机构:
Zenodo
创建时间:
2026-01-19
二维码
社区交流群
二维码
科研交流群
商业服务