BSTC (Baidu Speech Translation Corpus)
收藏资源简介:
BSTC是由百度公司创建的大规模中英双语语音翻译数据集,包含约68小时的普通话数据及其人工转录和英译文本,以及自动语音识别(ASR)模型的转录结果。该数据集旨在推动自动同声传译的研究和实用系统的发展,适用于自动同声传译系统的评估。数据集内容涵盖多个领域,如IT、经济、文化等,通过收集授权的视频讲座构建而成。
BSTC is a large-scale Chinese-English bilingual speech translation dataset created by Baidu. It includes approximately 68 hours of Mandarin speech data paired with their manual transcriptions, English translated texts, as well as the transcription results generated by automatic speech recognition (ASR) models. This dataset aims to advance research on automatic simultaneous interpretation and the development of practical application systems, and serves as a benchmark for evaluating automatic simultaneous interpretation systems. The dataset covers multiple domains such as IT, economics, culture and others, and is constructed by collecting authorized video lectures.




