err-video-news-transcribed
收藏资源简介:
Transcribed ERR Video News Dataset 是一个包含爱沙尼亚国家广播公司(ERR)视频新闻故事转录文本的数据集。该数据集包含约40,000个新闻故事,总时长约4000小时。转录文本是通过自动语音识别技术(gemini-3-flash-preview)生成的,并使用了上下文偏置技术以提高识别质量,平均词错误率(WER)约为5%。数据集经过严格过滤,仅保留视频中主要为爱沙尼亚语的故事,排除了大量非爱沙尼亚语或音乐的内容。每个新闻故事提供标题、导语文本、新闻正文以及视频新闻故事的转录文本。为避免版权问题,数据集中不包含音频或视频数据,但提供了每个故事原始视频的网页链接。该数据集适用于自动语音识别(ASR)任务及其他与爱沙尼亚语相关的自然语言处理研究。
The Transcribed ERR Video News Dataset is a corpus of transcribed texts for video news stories from the Estonian Public Broadcasting (ERR). It comprises roughly 40,000 news stories with a combined total duration of approximately 4,000 hours. The transcribed texts were generated via automatic speech recognition (ASR) using the gemini-3-flash-preview model, with contextual biasing implemented to enhance recognition performance; the average word error rate (WER) stands at around 5%. The dataset has been strictly filtered to retain only news stories predominantly in Estonian, excluding a significant portion of non-Estonian content or music-only segments. Each entry includes the news story's title, lead paragraph, main body, and the full transcribed text of the corresponding video news segment. To mitigate copyright concerns, the dataset does not include any raw audio or video files, but provides direct web links to the original videos for every news story. This dataset is suitable for automatic speech recognition (ASR) tasks and other Estonian-language focused natural language processing (NLP) research.
Transcribed ERR Video News Dataset 概述
数据集基本信息
- 许可证:cc-by-sa-4.0
- 任务类别:自动语音识别
- 语言:爱沙尼亚语 (et)
数据来源与内容
- 数据来源于爱沙尼亚国家广播公司 (https://www.err.ee/) 的视频新闻报道。
- 包含约 40,000 个新闻故事,总时长约 4000 小时。
- 转录文本通过语音识别系统 (gemini-3-flash-preview) 自动生成,并使用了基于同主题文本新闻故事的上下文偏置技术以提高自动语音识别质量。
- 转录的平均词错误率约为 5%。
- 数据集经过严格过滤,仅保留视频内容主要为爱沙尼亚语语音的新闻故事,已移除包含大量非爱沙尼亚语语音和/或音乐的新闻片段。
数据字段说明
每个新闻故事提供以下信息:
- 标题 (heading)
- 导语文本 (leadin text)
- 文本新闻故事正文 (main body of the textual news story)
- 视频新闻故事转录文本 (transcript of the video news story)
- 字幕 (subtitles):内容与“转录”字段相同,但以 VTT 格式分段,每个字幕块代表一个句子。句子/字幕的开始和结束时间通过强制对齐获得。
数据使用说明
- 为避免版权问题,数据集中不包含音频/视频数据。
- 为每个故事提供了可爬取原始视频的网页链接。
- 如需下载音频帮助,请联系 tanel.alumae@taltech.ee。




