LibriVoxDeEn - A Corpus for German-to-English Speech Translation and Speech Recognition

Name: LibriVoxDeEn - A Corpus for German-to-English Speech Translation and Speech Recognition
Creator: heiDATA
Published: 2025-01-28 12:49:43
License: 暂无描述

DataCite Commons2025-01-28 更新2025-04-17 收录

下载链接：

https://heidata.uni-heidelberg.de/citation?persistentId=doi:10.11588/DATA/TMEDTX

下载链接

链接失效反馈

官方服务：

资源简介：

This dataset is a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audio books. The corpus consists of over 100 hours of audio material and over 50k parallel sentences. The speech data are low in disfluencies because of the audio book setup. The quality of audio and sentence alignments has been checked by a manual evaluation, showing that that speech alignment is in general very high. The sentence alignment quality is comparable to well-used parallel translation data and can be adjusted by cutoffs on the automatic alignment score. To our knowledge, this corpus is to date the largest resource for end-to-end speech translation for German.

提供机构：

heiDATA

创建时间：

2019-10-21

5,000+

优质数据集

54 个

任务类型

进入经典数据集