遇见数据集

LiveLingo Voice Translation Benchmarks 2026

收藏
Zenodo2026-07-25 更新2026-08-02 收录
官方服务:

资源简介:

Benchmark data for real-time voice translation quality (speech in, translation out), covering both frontier speech-to-speech models and production speech-translation stacks. Headline results. On comprehension fidelity (0-5, n=120, English into Spanish / Simplified Chinese / Japanese / German): LiveLingo 4.96, Gemini Live 4.93, Google 4.77, Azure 4.65, Whisper + GPT-4o-mini 4.63, OpenAI gpt-realtime-translate 4.53. On median latency from end of speech to final translation: LiveLingo 1.5 s, Gemini Live 3.1 s, OpenAI gpt-realtime-translate 3.8 s, Azure 4.8 s, Google 26.7 s. LiveLingo also placed first on all 16 non-English corridors tested. Two studies. (1) A 16-language-corridor comprehension benchmark, with Turkish, Arabic, Indonesian, Vietnamese, Portuguese, Polish, Chinese and Malay source languages into the languages of their main migration destinations (tr-de, ar-de, ar-fr, ar-es, id-ja, id-ko, id-zh, vi-ko, vi-ja, pt-de, pt-fr, pl-de, zh-vi, zh-id, zh-ms, ms-zh). (2) A latency and comprehension-fidelity benchmark (n=120, English into Spanish / Simplified Chinese / Japanese / German), including per-utterance results for Gemini Live and the OpenAI Realtime API. Method. Every system receives identical source audio and runs its own complete speech-translation pipeline. Three independent LLM judges (GPT-4o, Gemini 2.5 Flash, Claude) score each translation 0-5 for comprehension fidelity, penalising wrong-word substitution, dropped entities and language-script mixing. Source audio is included so results can be reproduced. Published by LiveLingo. Live leaderboards: latency and comprehension benchmark and 16-corridor benchmark. Repository mirror: github.com/livelingo/voice-translation-benchmark.

提供机构:
Zenodo
创建时间:
2026-07-07
二维码
社区交流群
二维码
科研交流群
商业服务