nineninesix/diamond-benchmark
收藏资源简介:
Diamond Benchmark是一个包含750个真实降质语音录音的数据集,用于评估语音恢复模型的性能。这些录音是真实降质的语音,而非合成或通过管道重新降质的干净音频。每个录音都带有实际捕获中的损伤,如低比特率编解码压缩、窄带宽、背景噪声和削波。数据集提供了一个固定基准,用于测量语音恢复模型从真实降质中恢复干净、工作室质量音频的能力。每个录音都附有参考文本,因此可以在感知质量和内容保存两个轴上同时评分。数据集包括750个降质的.wav录音,语言为英语,结构包含音频文件和manifest.csv元数据文件,其中包含集合、ID、原始录音ID、音频路径、采样率、时长、说话者和文本等列。评分使用DNSMOS-P.835(感知质量)和CER(内容保存)两个互补指标。
Diamond Benchmark is a dataset containing 750 real degraded speech recordings for evaluating speech restoration models. These are real, genuinely degraded speech recordings — not synthetic, not clean audio re-degraded by a pipeline. Each clip carries the damage of an actual real-world capture: low-bitrate codec compression, narrow bandwidth, background noise, and clipping. Together they form a fixed benchmark for measuring how well a speech-restoration model recovers clean, studio-quality audio from real degradation. Every clip ships with its reference transcript, so restoration can be scored on two axes at once — perceptual quality and content preservation. The dataset includes 750 degraded .wav recordings in English, structured with audio files and a manifest.csv metadata file containing columns such as set, id, emolia_id, audio_path, sample_rate, duration_sec, speaker, and text. Scoring uses two complementary metrics: DNSMOS-P.835 (perceptual quality) and CER (content preservation).




