Trelis/multimed-test-filtered-v2-dropped
收藏资源简介:
--- language: - en tags: - whisper - filtered-dropped - speech - speech-to-text configs: - config_name: default data_files: - split: train path: data/train-* dataset_info: features: - name: audio dtype: audio: sampling_rate: 16000 - name: text dtype: string - name: text_ts dtype: string - name: preconditioning dtype: string - name: start_time dtype: string - name: end_time dtype: string - name: speech_duration dtype: float32 - name: word_timestamps dtype: string - name: source_file dtype: string - name: language dtype: string - name: filter_cer dtype: float64 - name: filter_model dtype: string - name: filter_confidence dtype: float64 splits: - name: train num_bytes: 718856722.7958573 num_examples: 2071 download_size: 721157697 dataset_size: 718856722.7958573 --- # Dropped Samples Samples removed during filtering of [None](https://huggingface.co/datasets/None). These had CER >= 47% against filter model fireworks/whisper-v3, indicating unreliable reference transcriptions. ## Columns | Column | Description | |--------|-------------| | `audio` | Audio sample | | `text` | Original reference transcription | | `filter_prediction` | Filter model prediction | | `filter_cer` | CER between filter prediction and reference | | `filter_model` | Model used for filter CER computation | | `filter_confidence` | Geometric mean word confidence (if available) | ## Related - **Kept dataset:** [Trelis/multimed-test-filtered-v2](https://huggingface.co/datasets/Trelis/multimed-test-filtered-v2) - **Source dataset:** [None](https://huggingface.co/datasets/None) --- *Generated by [Trelis Studio](https://studio.trelis.com)*
语言: - 英语 标签: - Whisper - 过滤剔除 - 语音 - 语音转文字 配置项: - 配置名称:默认配置 数据文件: - 划分集:训练集 路径:data/train-* 数据集信息: 特征: - 名称:audio 数据类型: 音频: 采样率:16000Hz - 名称:text 数据类型:字符串 - 名称:text_ts 数据类型:字符串 - 名称:preconditioning 数据类型:字符串 - 名称:start_time 数据类型:字符串 - 名称:end_time 数据类型:字符串 - 名称:speech_duration 数据类型:float32 - 名称:word_timestamps 数据类型:字符串 - 名称:source_file 数据类型:字符串 - 名称:language 数据类型:字符串 - 名称:filter_cer 数据类型:float64 - 名称:filter_model 数据类型:字符串 - 名称:filter_confidence 数据类型:float64 划分集信息: - 划分集名称:训练集 占用字节数:718856722.7958573 样本总数:2071 下载大小:721157697字节 数据集总大小:718856722.7958573字节 # 剔除样本 本数据集为[None](https://huggingface.co/datasets/None)数据集过滤流程中被移除的样本。这些样本与过滤模型fireworks/whisper-v3的字符错误率(Character Error Rate,CER)≥47%,表明其参考转录文本不可靠。 ## 字段说明 | 字段名 | 字段说明 | |--------|-------------| | `audio` | 音频样本 | | `text` | 原始参考转录文本 | | `filter_prediction` | 过滤模型的预测转录结果 | | `filter_cer` | 过滤模型预测结果与参考转录文本间的字符错误率 | | `filter_model` | 用于计算过滤字符错误率的模型 | | `filter_confidence` | 单词置信度的几何平均值(若可用) | ## 相关资源 - **保留样本数据集:** [Trelis/multimed-test-filtered-v2](https://huggingface.co/datasets/Trelis/multimed-test-filtered-v2) - **源数据集:** [None](https://huggingface.co/datasets/None) *由[Trelis Studio](https://studio.trelis.com)生成*




