CoSHE-500
收藏资源简介:
CoSHE-500(对话式印地英语语音评估)是一个专门用于评估自动语音识别系统在印地语-英语代码转换(Hinglish)场景下性能的数据集。该数据集包含500个对话语音片段,每个片段时长约30秒,采样率为16kHz单声道,并提供了逐字记录的混合脚本(天城文用于印地语,拉丁文用于英语)参考转录。数据内容为真实的对话语音,其中单个话语内存在印地语和英语的代码转换现象。该数据集旨在衡量ASR模型在现实世界代码转换环境中的性能,这是许多印地语/英语模型性能下降的典型场景。数据集仅用于评估和基准测试,不建议用于训练,以保持其作为公平公开基准的完整性。使用时应采用脚本安全的印度语文本规范化流程进行词错误率计算,并注意声明输出文本的书写规范。数据集基于CC-BY-NC-4.0许可分发,要求署名且仅限非商业用途。
CoSHE-500 (Conversational Speech in Hinglish for Evaluation - 500) is a dataset specifically designed for evaluating the performance of automatic speech recognition systems in Hindi-English code-switching (Hinglish) scenarios. It contains 500 conversational speech clips, each approximately 30 seconds long, with a sampling rate of 16kHz mono, and provides verbatim reference transcriptions in a mixed script (Devanagari for Hindi, Latin for English). The data consists of real conversational speech featuring intra-utterance code-switching between Hindi and English. The dataset aims to measure the quality of ASR models in real-world code-switching environments, a typical scenario where many Hindi/English models experience performance degradation. It is intended solely for evaluation and benchmarking, not for training, to preserve its integrity as a fair and open benchmark. When used, a script-safe Indian text normalization pipeline should be applied for word error rate calculation, and the writing standard of the output text should be declared. The dataset is distributed under the CC-BY-NC-4.0 license, requiring attribution and permitting only non-commercial use.




