遇见数据集

TaurenMountain/WenetSpeech-Formal

收藏
Hugging Face2026-05-20 更新2026-05-31 收录
官方服务:

资源简介:

WenetSpeech-Formal是一个基于WenetSpeech构建的中文语音数据集,增加了正式书面体转写。每个音频样本包含两个级别的转录:original_text(口语原文,含口语习惯表达)和target_text(正式书面体转写,由大语言模型辅助生成)。该数据集旨在训练将口语中文转换为正式书面中文的模型,例如用于自动语音识别后处理、口语到书面语规范化以及语音理解任务。数据集包含训练、验证和测试分割,分别来自WenetSpeech的L、DEV、TEST_NET和TEST_MEETING集合,总样本数超过100万。音频为16kHz单声道WAV格式,语言为中文。

WenetSpeech-Formal is a Chinese speech dataset built upon WenetSpeech, augmented with formal written-style transcriptions. Each audio sample contains two levels of transcription: original_text (original verbatim transcription with spoken habits) and target_text (formal written-style transcription, generated with the assistance of large language models). This dataset is designed for training models that convert spoken Chinese to formal written Chinese, such as ASR post-processing, spoken-to-written normalization, and speech understanding. It includes train, validation, and test splits derived from WenetSpeechs L, DEV, TEST_NET, and TEST_MEETING sets, with over 1 million samples total. Audio is in 16kHz mono WAV format, and the language is Chinese.

提供机构:
TaurenMountain
二维码
社区交流群
二维码
科研交流群
商业服务