Shriyaask/fleurs
收藏资源简介:
FLEURS数据集是FLoRes机器翻译基准的语音版本,用于评估跨语言、任务、领域和数据制度的语音表示。它涵盖了来自10多个语系的102种语言,3个不同领域和4个任务家族:语音识别、翻译、分类和检索。数据集使用FLoRes开发和开发测试公开集中的2009个n-way平行句子,每种语言的训练集大约有10小时的监督数据。训练集的说话者与开发/测试集的说话者不同。采用多语言微调,并平均所有语言的“单位错误率”(字符、符号)。语言和结果还分为七个地理区域:西欧、东欧、中亚/中东/北非、撒哈拉以南非洲、南亚、东南亚和CJK语言。
Fleurs is the speech version of the FLoRes machine translation benchmark, designed to evaluate speech representations across languages, tasks, domains, and data regimes. It covers 102 languages from over 10 language families, 3 different domains, and 4 task families: speech recognition, translation, classification, and retrieval. The dataset uses 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, with around 10 hours of supervision per language in the training sets. Speakers of the train sets are different from those in the dev/test sets. Multilingual fine-tuning is used, and the unit error rate (characters, signs) of all languages is averaged. Languages and results are also grouped into seven geographical areas: Western Europe, Eastern Europe, Central-Asia/Middle-East/North-Africa, Sub-Saharan Africa, South-Asia, South-East Asia, and CJK languages.




