MockConf
收藏资源简介:
MockConf 是一个学生口译数据集,由模拟会议中收集的学生口译记录组成。该数据集包含7小时、5种欧洲语言的录音,并已转录和按单词和跨度级别进行了对齐。数据集创建过程涉及从模拟会议中获取忠实的人类口译转录,然后使用InterAlign工具手动对齐和注释。MockConf数据集和InterAlign工具已公开发布,可用于语言分析、自动对齐工具的开发和评估、教育目的以及自动同声传译系统的评估。
MockConf is a student interpreting dataset composed of student interpreting recordings collected from simulated conferences. The dataset contains 7 hours of recordings in 5 European languages, and has been transcribed and aligned at both word and span levels. The dataset creation process involved obtaining faithful human interpreting transcripts from simulated conferences, followed by manual alignment and annotation using the InterAlign tool. The MockConf dataset and the InterAlign tool have been publicly released, and can be used for linguistic analysis, the development and evaluation of automatic alignment tools, educational purposes, as well as the evaluation of automatic simultaneous interpreting systems.
MockConf数据集概述
数据集基本信息
- 数据集名称:MockConf
数据集描述
(根据提供的README内容,该数据集未包含具体描述信息)




