ASR-EC
收藏资源简介:
ASR-EC数据集是由香港科技大学计算机科学与工程系创建的,专门用于评估大型语言模型在中文自动语音识别错误纠正方面的能力。该数据集包含来自THCHS-30、AISHELL-1、AISHELL-2和WeNetSpeech的音频数据,总计约544,551条语音记录。数据集的创建过程基于Kaldi-K1和Kaldi-K2工具,通过处理这些音频数据生成错误转录,以模拟实际应用中的语音识别错误。ASR-EC数据集主要应用于语音识别系统的错误纠正研究,旨在提高语音识别系统的准确性和鲁棒性。
The ASR-EC Dataset was developed by the Department of Computer Science and Engineering at the Hong Kong University of Science and Technology, specifically designed to evaluate the performance of large language models (LLMs) in Chinese automatic speech recognition (ASR) error correction. This dataset includes audio data sourced from THCHS-30, AISHELL-1, AISHELL-2 and WeNetSpeech, totaling approximately 544,551 speech recordings. The dataset was constructed using the Kaldi-K1 and Kaldi-K2 toolkits, where erroneous transcriptions were generated by processing the aforementioned audio data to simulate realistic speech recognition errors in real-world applications. The ASR-EC dataset is primarily utilized for research on error correction in speech recognition systems, with the objective of improving the accuracy and robustness of speech recognition systems.

- 1ASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction香港科技大学计算机科学与工程系 · 2024年



