遇见数据集

channelcorp/meet_llasa_1_vad_asr_clean_esther_cer_added_end_refine

收藏
Hugging Face2025-12-17 更新2025-12-20 收录
官方服务:

资源简介:

该数据集包含音频文件及其相关元数据,如时间戳、持续时间、描述文本、说话者信息和预测文本。时间戳结构详细记录了音频片段的开始、结束时间以及对应的文本内容。此外,数据集还提供了字符错误率(CER)和精细化的结束时间信息,包括调试信息和截断原因等。数据集仅包含训练集,共有14,407个示例,总大小为7,797,714字节。

This dataset contains audio files and related metadata, such as timestamps, duration, description text, speaker information, and predicted text. The timestamp structure records the start and end times of audio segments and the corresponding text content in detail. Additionally, the dataset provides character error rate (CER) and refined end-time information, including debugging details and truncation reasons. The dataset only includes a training set with 14,407 examples and a total size of 7,797,714 bytes.

提供机构:
channelcorp
二维码
社区交流群
二维码
科研交流群
商业服务