CS-Dialogue
收藏资源简介:
CS-Dialogue是一个大规模的、公开可用的普通话-英语代码转换语音对话数据集。该数据集解决了现有代码转换语音数据集中存在的主要问题,如数据集规模小、缺乏自然对话和缺失完整对话录音。它为推进代码转换自动语音识别(ASR)和其他相关领域的研究提供了坚实的基础。数据集包含104.02小时的自发对话录音,由200位讲者录制的100对两个人的对话组成。数据集在CC BY-NC-SA 4.0许可下发布,意味着它可用于非商业用途。
CS-Dialogue is a large-scale, publicly available Mandarin-English code-switching speech dialogue dataset. It addresses the core limitations of existing code-switching speech datasets, including small dataset size, lack of natural dialogues, and missing complete conversation recordings. This dataset provides a solid foundation for advancing research in code-switching automatic speech recognition (ASR) and other related domains. It contains 104.02 hours of spontaneous conversational speech recordings, comprising 100 two-person dialogues recorded by 200 total speakers. The dataset is released under the CC BY-NC-SA 4.0 license, which permits non-commercial use.




