NbAiLab/NPSC
收藏资源简介:
挪威议会语音语料库(NPSC)是由挪威国家图书馆的挪威语言银行在2019-2021年间创建的语音语料库。该语料库包含挪威议会会议的录音和对应的挪威书面转录,转录工作由训练有素的语言学家完成,并经过校对以确保准确性和一致性。数据集适用于自动语音识别和音频分类任务。
The Norwegian Parliament Speech Corpus (NPSC) is a speech corpus created by the Norwegian Language Bank of the National Library of Norway between 2019 and 2021. This corpus contains audio recordings of Norwegian parliamentary sessions and their corresponding Norwegian written transcripts, which were completed by trained linguists and proofread to ensure accuracy and consistency. This dataset is applicable to automatic speech recognition and audio classification tasks.
数据集概述
数据集名称
- 名称: NPSC (Norwegian Parliamentary Speech Corpus)
语言信息
- 语言:
- 主要语言: nb (Norwegian Bokmål), nn (Norwegian Nynorsk)
- 辅助语言: no, en-US
许可证
- 许可证: CC0-1.0
数据集大小
- 大小: 2G<n<1B
数据来源
- 来源: 原始数据
任务类别
- 任务类别:
- automatic-speech-recognition
- audio-classification
数据集标签
- 标签: speech-modeling
数据集描述
- 描述: NPSC是由挪威语言银行在2019-2021年间创建的语音语料库,包含挪威议会的演讲录音及其对应的挪威语Bokmål和Nynorsk的正字法转录。所有转录由训练有素的语言学家或语文学家手动完成,并经过校对以确保一致性和准确性。
数据字段
- 数据字段:
- sentence_id: 句子唯一标识符
- sentence_order: 句子在会议中的顺序
- speaker_id: 发言人ID
- meeting_date: 会议日期
- speaker_name: 发言人姓名
- sentence_text: 句子文本
- sentence_language_code: 句子语言代码
- text: 句子文本副本
- start_time: 句子开始时间
- end_time: 句子结束时间
- normsentence_text: 规范化句子文本
- transsentence_text: 翻译后的句子文本
- translated: 翻译标识
- audio: 音频数据
数据集统计
- 统计信息:
- 总时长(含停顿): 140.3小时
- 总时长(不含停顿): 125.7小时
- 单词计数: 120万
- 句子计数: 64,531
- 语言分布: Nynorsk 12.8%, Bokmål 87.2%
- 性别分布: 女性38.3%, 男性61.7%
许可证信息
- 音频和转录: CC0-1.0
- HuggingFace数据集整理: CC-BY-SA-3.0
引用信息
-
引用:
@inproceedings{solberg2022norwegian, title={The Norwegian Parliamentary Speech Corpus}, author={Solberg, Per Erik and Ortiz, Pablo}, booktitle={Proceedings of the 13th Language Resources and Evaluation Conference}, url={http://www.lrec-conf.org/proceedings/lrec2022/pdf/2022.lrec-1.106.pdf}, year={2022} }




