AgentPublic/eval-stt-officiels
收藏资源简介:
EvalSTT — 官方语料库(法语)是一个公共评估语料库,用于评估语音转文本模型在法国行政语言上的表现,包括官方演讲、公开讲话和政府质询。该语料库由法国国家数字事务局(DINUM)的国家人工智能部门构建,旨在评估语音转文本模型。数据集发布是为了透明性:它记录了用于评估的数据集,并允许复现测量结果(如词错误率WER和语义严重性)。内容包括音频文件和参考转录,共13条记录:5条短演讲(≤10分钟)、5条长演讲(15-39分钟)和3条政府质询会议。方法学警告指出参考转录是平滑的(基于机构字幕或笔录),已清理不流畅部分,因此非严格逐字记录,这会导致绝对WER偏高,但所有比较模型的偏差相同。数据集仅包含公共数据,音频和参考来源自官方公共资源(如info.gouv.fr、SIG、国民议会、参议院、爱丽舍宫),使用Etalab 2.0开放许可证。
EvalSTT — Corpus officiels (FR) is a public evaluation corpus for speech-to-text models on French administrative language: official speeches, public addresses, and government questions. It is constituted by the AI in the State department (DINUM) as part of the evaluation of speech-to-text models. This dataset is published for transparency: it documents the datasets used for our evaluations and allows reproducing measurements (WER, semantic severity). Content includes audio files and reference transcriptions, with 13 recordings: 5 short speeches (≤10 min), 5 long-form speeches (15–39 min), and 3 government question sessions. Methodological caution notes that references are smoothed (institutional subtitles/transcripts) with disfluencies cleaned, thus not strictly verbatim, which raises absolute WER but bias is identical for all compared models. The dataset contains only public data, sourced from official public sources (info.gouv.fr, SIG, National Assembly, Senate, Élysée) under Etalab 2.0 license.




