scribe-project/nbtale3
收藏资源简介:
这是NB Tale模块3的博克马尔语片段版本,用于测试论文《Improving Generalization of Norwegian ASR with Limited Linguistic Resources》中的模型,该论文在NoDaLiDa 2023会议上发表。数据集仅包含长度小于15秒的片段,且包含本地和非本地说话者的语音片段。在论文分析中,`region`字段为`foreign`的说话者被过滤掉了。数据集的特征包括说话者ID、性别、话语ID、语言、原始文本、完整音频文件、原始数据分割、地区、持续时间、开始时间、结束时间、话语音频文件和标准化文本。数据集分为训练集,包含8033个样本,总大小为1233495883.99字节。数据集的许可证为CC0。
This is the segmented Bokmål version of NB Tale Module 3, used for testing models in the paper *Improving Generalization of Norwegian ASR with Limited Linguistic Resources*, which was published at the NoDaLiDa 2023 conference. This dataset only contains segments with a duration of less than 15 seconds, and includes speech segments from both native and non-native speakers. In the paper's analysis, speakers with the `region` field marked as `foreign` were filtered out. The features of this dataset include Speaker ID, Gender, Utterance ID, Language, Raw Text, Full Audio File, Original Data Split, Region, Duration, Start Time, End Time, Utterance Audio File, and Normalized Text. This dataset is split into the training set, which contains 8033 samples with a total size of 1233495883.99 bytes. The license for this dataset is CC0.
数据集概述
数据集特征
- speaker_id: 字符串类型
- gender: 字符串类型
- utterance_id: 字符串类型
- language: 字符串类型
- raw_text: 字符串类型
- full_audio_file: 字符串类型
- original_data_split: 字符串类型
- region: 字符串类型
- duration: 浮点数类型
- start: 浮点数类型
- end: 浮点数类型
- utterance_audio_file: 音频类型
- standardized_text: 字符串类型
数据集分割
- train:
- num_bytes: 1233495883.99
- num_examples: 8033
数据集大小
- download_size: 1287266972
- dataset_size: 1233495883.99
数据集描述
- 包含时长小于15秒的挪威语Bokmål段落。
- 数据集用于测试论文中提到的模型,该论文讨论了在有限语言资源下提高挪威自动语音识别的一般化能力。
- 数据集包含本地和非本地说话者,但在论文分析中排除了地区设置为
foreign的说话者。
语言
- 挪威语 Bokmål
数据集创建
- 源数据: 来自挪威语言银行的完整数据集。
- 初始数据收集和规范化: 使用Spraakbanken下载器获取数据,并通过结合数据集标准化脚本进行标准化。
许可证信息
- CC0
引用信息
@inproceedings{ solberg2023improving, title={Improving Generalization of Norwegian {ASR} with Limited Linguistic Resources}, author={Per Erik Solberg and Pablo Ortiz and Phoebe Parsons and Torbj{o}rn Svendsen and Giampiero Salvi}, booktitle={The 24rd Nordic Conference on Computational Linguistics}, year={2023} }




