espnet/yodas-granary
收藏资源简介:
YODAS-Granary是一个高质量伪标记语音数据集,专注于自动语音识别(ASR)和自动语音翻译(AST)任务,涵盖了23种欧洲语言。数据集由NVIDIA/Granary数据集的子集组成,并提供两种任务的核心数据:ASR和AST。ASR数据包含23种欧洲语言的伪标记转录,而AST数据包含22种非英语语言的英语高质量翻译。数据集使用Systran/faster-whisper-large-v3模型进行转录,并使用Qwen/Qwen2.5-7B-Instruct模型进行翻译。
YODAS-Granary is a curated subset of the larger `nvidia/Granary` dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. The dataset is derived from the `espnet/yodas2` corpus and provides high-quality pseudo-labeled speech data. The ASR data covers 23 European languages with pseudo-labeled transcriptions generated using the `Systran/faster-whisper-large-v3` model and post-processed to restore punctuation and capitalization using `Qwen/Qwen2.5-7B-Instruct`. The AST data covers 22 non-English languages and consists of high-quality translations into English generated from the ASR subset using the `utter-project/EuroLLM-9B-Instruct` model.




