ArtificialAnalysis/Earnings22-Cleaned-AA-chunked
收藏资源简介:
Earnings22-Cleaned-AA-chunked 是一个用于流式语音转文本评估的音频数据集,基于Earnings22-Cleaned-AA数据集的块状版本。原始数据来源于esb/datasets的企业财报电话会议语料库,该子集经过人工清理和校正参考转录文本,以减少地面真实错误对词错误率评估的不公平影响。数据集包含6个父样本,分割成341个音频块,每个块时长在2.32至29.98秒之间,总时长约115分钟,语言为英语。音频块使用Silero VAD进行语音活动检测生成,相邻片段被组合成不超过30秒的块,以支持流式语音转文本系统。数据集提供每个块的元数据(如ID、起始时间、文件名),但不包含每块的参考转录文本;评分时参考文本从清理后的父转录文本中派生并在评估流程中对齐。
Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original data comes from esb/datasets, a corpus of corporate earnings calls, with manual review and correction of reference transcripts to reduce ground-truth errors affecting word error rate. It contains 6 parent samples split into 341 audio chunks, each ranging from 2.32 to 29.98 seconds in duration, totaling approximately 115 minutes, in English. Chunks are generated using Silero VAD to identify speech segments, combined into segments up to 30 seconds to support streaming STT systems. The dataset includes metadata per chunk (e.g., ID, timestamps, filename) but not per-chunk reference transcripts; reference text is derived from cleaned parent transcripts and aligned during evaluation.




