eval-stt-officiels
收藏资源简介:
EvalSTT — Corpus officiels (FR) 是一个用于评估语音转文本(ASR)模型在法国行政语言上的公开数据集,由法国国家数字化与信息技术部(DINUM)的“国家人工智能”部门构建。该数据集旨在提高模型评估的透明度,允许复现词错误率(WER)和语义严重性等度量。数据集内容聚焦于官方行政语言,包括官方演讲、公开讲话和政府质询会议。具体包含8个录音文件(每个时长约11分钟以内),分为两类:5个短篇官方演讲(时长≤10分钟),其参考转录为“平滑”版本,即清理了不流利部分的机构性逐字稿或字幕;3个政府质询会议录音(约10分钟,来自法国国民议会),其参考转录为“严格逐字”版本,保留了说话中的不流利、重复和错误起始。数据以音频文件(.mp3格式)和对应的参考转录文本文件(.txt格式)提供。需要注意的是,数据集中并存两种不同的参考转录制度(平滑与严格逐字),因此不应在绝对意义上比较来自这两类数据的WER,也不应与使用不同转录制度的第三方语料库进行直接比较。所有数据均来源于官方公共渠道(如info.gouv.fr、SIG、国民议会、参议院、爱丽舍宫),并采用Etalab 2.0开放许可证发布,仅包含公共领域数据。
EvalSTT — Corpus officiels (FR) is a public dataset designed for evaluating speech-to-text (ASR) models on French administrative language, constructed by the National Artificial Intelligence department of the French Directorate for Digital and Information Technology (DINUM). The dataset aims to enhance transparency in model evaluation, allowing for the reproduction of metrics such as Word Error Rate (WER) and semantic severity. It focuses on official administrative language, including official speeches, public addresses, and government inquiry sessions. Specifically, it contains 8 audio files (each under approximately 11 minutes), categorized into two types: 5 short official speeches (duration ≤ 10 minutes) with reference transcriptions in a smoothed version, which are institutional verbatim transcripts or subtitles cleaned of disfluencies; and 3 recordings of government inquiry sessions (about 10 minutes, from the French National Assembly) with reference transcriptions in a strict verbatim version, preserving disfluencies, repetitions, and false starts in speech. The data is provided as audio files (.mp3 format) and corresponding reference transcription text files (.txt format). It is important to note that the dataset includes two different reference transcription regimes (smoothed and strict verbatim), so WER from these two data types should not be compared in absolute terms, nor should direct comparisons be made with third-party corpora using different transcription regimes. All data is sourced from official public channels (e.g., info.gouv.fr, SIG, National Assembly, Senate, Élysée Palace) and released under the Etalab 2.0 open license, containing only public domain data.
数据集概述
EvalSTT — Corpus officiels (FR) 是一个用于评估法语语音转文字(speech-to-text)模型性能的公共数据集,专注于法国行政语言,包括官方演讲、公开讲话和政府质询会议。由法国政府人工智能部门(DINUM) 构建,主要用于模型评估的透明性和可复现性。
数据集规模与格式
- 规模:样本数量小于1000(具体为8个录音文件)。
- 格式:
audio/<id>.mp3:音频源文件(公共演讲/会议)。ground_truth/<id>.txt:对应的参考转录文本。
- 时长:每个录音不超过约11分钟,时长均匀。
数据内容
- 5个官方简短演讲(≤10分钟):参考转录为平滑版(即清理了不流畅部分,如副标题或官方机构提供的规范文本)。
- 3个政府质询会议(约10分钟,来自法国国民议会):参考转录为严格逐字版(保留口误、重复、虚假开头等口语特征)。
重要方法提示
数据集中存在两种不同的参考转录标准,绝对WER(词错误率)值不能在不同标准间直接比较:
- 官方演讲(5个):使用平滑版参考(非严格逐字),WER绝对值会偏高,但所有模型在此标准下偏差一致。
- 政府质询(3个):使用严格逐字版参考(基于实际音频重建,而非国民议会的正式记录),保留所有口语特征。
数据范围与来源
- 数据范围:仅包含公共领域的数据。
- 来源:音频和参考转录来自官方公共来源(info.gouv.fr、SIG、法国国民议会、参议院、爱丽舍宫)。
- 许可证:Etalab 2.0(法国开放许可证)。





