sarvam-dub-benchmark-set
收藏资源简介:
language: - en - hi - bn - ta - te - kn - ml - mr - gu - or - pa license: other task_categories: - text-to-speech - audio-to-audio pretty_name: Sarvam Dubbing Benchmark Dataset size_categories: - 1K<n<10K --- # Sarvam Dubbing Benchmark Dataset ## Dataset Description Multilingual evaluation dataset for real-time dubbing and voice cloning benchmarking with focus on speaker similarity preservation across same-lingual and cross-lingual scenarios. This dataset was used to benchmark production dubbing systems. Internal evaluations showed higher speaker similarity than ElevenLabs v3 and Cartesia Sonic under an identical scoring protocol. Check out the Sarvam Dub blog for more details: https://www.sarvam.ai/blogs/sarvam-dub **Evaluation-only dataset — not for training.** ## Supported Languages This benchmark includes the following languages: - English (en) - Hindi (hi) - Bengali (bn) - Tamil (ta) - Telugu (te) - Kannada (kn) - Malayalam (ml) - Marathi (mr) - Gujarati (gu) - Odia (or) - Punjabi (pa) ## Dataset Summary - Speakers: 64 - Languages per speaker: 11 - Total samples: 704 - Setup: One-shot speaker conditioning - Metric: Speaker similarity ## Dataset Schema Each record contains: - `reference_audio` — speaker prompt audio - `target_text` — text to dub - `target_language` — output language code ## Speaker Similarity Scoring Speaker similarity is computed using the SpeechBrain ECAPA speaker embedding model: Model: https://huggingface.co/speechbrain/spkrec-ecapa-voxceleb ## Intended Use Dubbing benchmarking, voice cloning evaluation, speaker similarity measurement.



