DEBATE
收藏资源简介:
DEBATE是一个专门用于研究汉语语音消歧的语音-文本数据集,它旨在探讨语音线索如何帮助解决文本歧义,并揭示说话者的真正意图。该数据集包含1001个精心挑选的歧义性语句,每个语句由10位母语者录音,共超过10K条音频记录。数据集涵盖了多音字歧义、结构歧义和焦点歧义三种类型。DEBATE数据集的创建过程包括收集原始文本、语音数据录制和质量控制三个主要阶段。该数据集可用于评估和训练模型在语音消歧方面的能力,并为构建类似的语音消歧数据集奠定了基础。
DEBATE is a speech-text dataset specifically developed for research on Chinese speech disambiguation. It aims to investigate how speech cues assist in resolving textual ambiguities and revealing speakers' genuine intentions. This dataset comprises 1001 carefully curated ambiguous utterances, each recorded by 10 native speakers, yielding more than 10,000 audio recordings. It covers three types of ambiguities: polyphonic character ambiguity, structural ambiguity, and focus ambiguity. The development of the DEBATE dataset includes three main stages: collecting original texts, recording speech data, and conducting quality control. This dataset can be used to evaluate and train models' speech disambiguation capabilities, and lay a foundation for building similar speech disambiguation datasets.
DEBATE数据集概述
数据集简介
- 名称:DEBATE (A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech)
- 类型:中文语音-文本数据集
- 目的:研究语音线索(发音、停顿、重音和语调)如何帮助解决文本歧义并揭示说话者真实意图
- 内容:包含1,001个精心挑选的歧义语句,每个语句由10名母语者录制
数据集统计
| 任务类型 | 样本数 | 时长(小时) | 平均时长(秒) | 时长范围(秒) |
|---|---|---|---|---|
| Task_Proun | 2,000 | 1.64 | 2.94 | 1.15-5.80 |
| Task_Pause | 4,010 | 4.28 | 3.84 | 1.60-11.80 |
| Task_Stres | 4,000 | 3.74 | 3.37 | 1.43-8.51 |
| 总计 | 10,010 | 9.66 | 3.47 | 1.15-11.80 |
数据来源与构建
- 文本来源:
- 开源语料库
- 社交媒体平台
- 标准化考试题库
- 歧义类型标注:
- 多音字歧义
- 结构歧义
- 焦点歧义
- 额外标注:
- 每个句子的语义注释
数据采集与处理
- 录制方式:10名人口特征平衡的说话者,使用手机双人协作录制
- 质量控制:
- 音频文件数量验证
- 随机抽样检查
- 使用ASR模型进行音频CER测试
- 后期处理:所有音频文件重采样至16kHz
数据结构
DEBATE_Audio
- speaker_x
- polyphone
- sx_p_id_000.wav
- sx_p_id_001.wav
- segment
- sx_seg_id_000.wav
- sx_seg_id_001.wav
- stress
- sx_s_id_000.wav
- sx_s_id_001.wav
- polyphone
标注文件
- Task_Proun.xls:包含转录句子、音频ID、语义注释和多音字具体发音
- Task_Pause.xls:包含转录句子、音频ID和语义注释
- Task_Stres.xls:包含转录句子、音频ID和语义注释
- metadata.xls:包含录音志愿者的ID、年龄、性别和地区信息
使用许可
- 许可证:CC BY-NC-4.0(仅限非商业用途)
- 获取方式:通过Zenodo获取(https://zenodo.org/records/15609922)
引用格式
bibtex @misc{guo2025debatedatasetdisentanglingtextual, title={DEBATE: A Dataset for Disentangling Textual Ambiguity in Mandarin Through Speech}, author={Haotian Guo and Jing Han and Yongfeng Tu and Shihao Gao and Shengfan Shen and Wulong Xiang and Weihao Gan and Zixing Zhang}, year={2025}, eprint={2506.07502}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2506.07502}, }
联系方式
- 反馈邮箱:haotianguo@hnu.edu.cn




