遇见数据集

AI-SPEAK: Serbian Audio-Visual Facial Animation Dataset

收藏
Zenodo2026-05-20 更新2026-05-26 收录
官方服务:

资源简介:

AI-SPEAK is a multimodal dataset specifically designed for speech-driven facial animation in the Serbian language. It contains synchronized audio (16 kHz) and 3D facial motion data (52 ARKit blendshape coefficients) captured at 60 fps. Dataset Statistics Language: Serbian (South Slavic) Speakers: 5 native speakers Total Duration: ~76 minutes (4,580 seconds) Total Utterances: 808 Data Formats: .wav (audio), .csv (blendshapes), .xlsx (transcripts), .txt (alignments) Phoneme Inventory 30 Serbian phonemes + silence token (sil), total 31 labels:a, b, v, g, d, đ, e, ž, z, i, j, k, l, lj, m, n, nj, o, p, r, s, t, ć, u, f, h, c, č, dž, š, sil Train/Validation Split The dataset is split into training (626 utterances) and validation (182 utterances). For the 30 shared sentences recorded by all speakers, the first 20 per speaker are assigned to training and the remaining 10 to validation.

提供机构:
Zenodo
创建时间:
2026-05-20
二维码
社区交流群
二维码
科研交流群
商业服务