AI-SPEAK: Serbian Audio-Visual Facial Animation Dataset
收藏资源简介:
AI-SPEAK is a multimodal dataset specifically designed for speech-driven facial animation in the Serbian language. It contains synchronized audio (16 kHz) and 3D facial motion data (52 ARKit blendshape coefficients) captured at 60 fps. Dataset Statistics Language: Serbian (South Slavic) Speakers: 5 native speakers Total Duration: ~76 minutes (4,580 seconds) Total Utterances: 808 Data Formats: .wav (audio), .csv (blendshapes), .xlsx (transcripts), .txt (alignments) Phoneme Inventory 30 Serbian phonemes + silence token (sil), total 31 labels:a, b, v, g, d, đ, e, ž, z, i, j, k, l, lj, m, n, nj, o, p, r, s, t, ć, u, f, h, c, č, dž, š, sil Train/Validation Split The dataset is split into training (626 utterances) and validation (182 utterances). For the 30 shared sentences recorded by all speakers, the first 20 per speaker are assigned to training and the remaining 10 to validation.



