toloka/CrowdSpeech
收藏资源简介:
CrowdSpeech是第一个公开的大规模众包音频转录数据集,基于LibriSpeech语料库在Toloka众包平台上构建。数据集包含22K个实例,约155K个众包转录注释。每个数据实例包含音频记录的URL、转录列表、执行者标识符和真实转录。数据集支持的任务包括众包转录的聚合,语言为英语。数据集结构包括五个分割:train、test、test.other、dev.clean和dev.other,分别对应LibriSpeech的clean和other部分。数据集的创建过程包括在Toloka平台上进行注释,注释者需要通过入门考试,并且每个任务由7个注释者完成。
CrowdSpeech is the first publicly available large-scale crowdsourced audio transcription dataset, constructed based on the LibriSpeech corpus via the Toloka crowdsourcing platform. The dataset contains 22K instances and approximately 155K crowdsourced transcription annotations. Each data instance includes the URL of the audio recording, a list of transcriptions, performer identifiers, and the ground-truth transcription. The supported tasks include crowdsourced transcription aggregation, with the target language being English. The dataset is split into five subsets: train, test, test.other, dev.clean, and dev.other, which respectively correspond to the clean and other partitions of LibriSpeech. The dataset creation process involves annotation work on the Toloka platform, where annotators must pass a qualifying exam, and each transcription task is completed by 7 annotators.
CrowdSpeech数据集概述
数据集描述
- 名称: CrowdSpeech
- 语言: 英语(en-US)
- 许可证: cc-by-4.0
- 数据来源: 原始数据,通过Toloka crowdsourcing平台对LibriSpeech数据集进行标注
- 数据规模: 22K实例,约155K注释
- 任务类别: 摘要生成、自动语音识别、文本到文本生成
- 标签: 条件文本生成、结构化到文本、语音识别
数据集结构
数据实例
- 内容: 每个实例包含音频录音的URL、一组转录文本及其对应的执行者标识和标准答案。
- 转录数量: 每个实例提供7个众包转录。
数据字段
- task: 音频录音的URL
- transcriptions: 众包转录文本列表
- performers: 执行者标识
- gt: 标准转录文本
数据分割
- 分割: 训练集、测试集、测试其他集、开发清洁集、开发其他集
- 特点: 训练集、测试集和开发清洁集对应LibriSpeech的高质量部分,开发其他集和测试其他集对应更具挑战性的部分。
数据集创建
源数据
- 来源: LibriSpeech,约1000小时的16kHz英语朗读音频
注释
- 平台: Toloka crowdsourcing平台
- 注释者筛选: 通过入口考试,要求Word Error Rate (WER) 不超过40%
- 注释重叠: 每个任务由7个注释者完成
引用信息
@inproceedings{CrowdSpeech, author = {Pavlichenko, Nikita and Stelmakh, Ivan and Ustalov, Dmitry}, title = {{CrowdSpeech and Vox~DIY: Benchmark Dataset for Crowdsourced Audio Transcription}}, year = {2021}, booktitle = {Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks}, eprint = {2107.01091}, eprinttype = {arxiv}, eprintclass = {cs.SD}, url = {https://openreview.net/forum?id=3_hgF1NAXU7}, language = {english}, pubstate = {forthcoming}, }




