johnatanebonilla/coser
收藏资源简介:
--- dataset_info: features: - name: audio dtype: audio - name: filename dtype: string - name: turno_id dtype: int64 - name: turno_time dtype: string - name: sentence dtype: string - name: sentence_fono dtype: string - name: sentence_fono_sin_marcas dtype: string - name: sentence_orto dtype: string - name: sentence_orto_sin_marcas dtype: string - name: Provincia dtype: string - name: Enclave dtype: string - name: Fecha dtype: string - name: Duración dtype: string - name: Informantes dtype: string splits: - name: train num_bytes: 4600923777.433 num_examples: 53971 - name: validation num_bytes: 503026194.46 num_examples: 6689 - name: test num_bytes: 486076659.954 num_examples: 6726 download_size: 4707509912 dataset_size: 5590026631.847 configs: - config_name: default data_files: - split: train path: data/train-* - split: validation path: data/validation-* - split: test path: data/test-* task_categories: - automatic-speech-recognition - conversational language: - es pretty_name: COSER-ASR Subset size_categories: - 10K<n<100K --- # Introduction The "COSER-ASR" Subset is a specialized extract from the "Corpus Oral y Sonoro del Español Rural" (COSER; Fernández-Ordóñez 2005-present), meaning the "Audible Corpus of Spoken Rural Spanish". This dataset has been specifically curated to facilitate the fine-tuning of Whisper, an automatic speech recognition system. For this purpose, audio and text segments ranging from 3 to 30 seconds have been automatically extracted from the COSER corpus. These segments provide concise and diverse samples of spoken rural Spanish, ideal for training and refining speech recognition models. To ensure manageability and efficient processing, a maximum of 1024 tokens were used in the dataset, striking a balance between comprehensive coverage and computational efficiency. # Content and Demographic Focus The original COSER dataset includes 218 transcriptions of semi-structured interviews primarily with elderly, less-educated individuals from rural Spain. These interviews, each averaging around 54 minutes, are rich in dialectal variations and linguistic nuances, offering valuable insights into traditional Spanish dialects. # Transcription Approach The "coser" dataset provides multiple layers of transcription to cater to different linguistic and computational needs: ### Original Transcription (sentence): This is the direct transcription of the audio segments, preserving the original speech as closely as possible and the complete original transcription. ### Phonological Approximation (sentence_fono): Here, the transcription is modified to reflect the phonological characteristics of the dialectal pronunciation. This version is crucial for understanding the phonetic nuances of rural Spanish dialects. ### Phonological Transcription without Discourse Markers (sentence_fono_sin_marcas): This transcription removes discourse markers such as laughter, assent, etc., that are typically enclosed in square brackets. It offers a cleaner version focusing solely on the spoken words. ### Orthographic Correspondence (sentence_orto): This layer provides the standard orthographic equivalent of the words transcribed phonologically. It bridges the gap between dialectal speech and standard Spanish orthography. ### Orthographic Transcription without Discourse Markers (sentence_orto_sin_marcas): Similar to the phonological version without markers, this transcription provides a standard orthographic text devoid of any discourse markers. This is particularly useful for applications requiring clean text data. # Limitations Limitations of this model include the fact that the time intervals in the COSER corpus are not systematically aligned, meaning that there may not be a perfect one-to-one correspondence between the audio and text data. # Additional Information and Resources To explore more about the COSER corpus, its methodologies, and the full range of transcriptions, visit http://coser.lllf.uam.es/ and http://coser.lllf.uam.es/transcripcion.php. These resources provide an in-depth look at the COSER project, detailing its comprehensive approach to capturing the linguistic diversity of rural Spanish. # References Fernández-Ordóñez, I. (Ed.). (2005-present). Corpus Oral y Sonoro del Español Rural. Retrieved April 15, 2022, from http://www.corpusrural.es/
dataset_info: 特征字段: - 字段名: 音频, 数据类型: 音频 - 字段名: 文件名, 数据类型: 字符串 - 字段名: 轮次ID, 数据类型: 64位整数 - 字段名: 轮次时间, 数据类型: 字符串 - 字段名: 语句, 数据类型: 字符串 - 字段名: 语音近似转写语句(sentence_fono), 数据类型: 字符串 - 字段名: 无标记语音转写语句(sentence_fono_sin_marcas), 数据类型: 字符串 - 字段名: 标准正写法对应语句(sentence_orto), 数据类型: 字符串 - 字段名: 无标记标准正写法转写语句(sentence_orto_sin_marcas), 数据类型: 字符串 - 字段名: 省份, 数据类型: 字符串 - 字段名: 区域, 数据类型: 字符串 - 字段名: 录制日期, 数据类型: 字符串 - 字段名: 时长, 数据类型: 字符串 - 字段名: 受访人信息, 数据类型: 字符串 划分集: - 划分名称: 训练集, 字节数: 4600923777.433, 样本量: 53971 - 划分名称: 验证集, 字节数: 503026194.46, 样本量: 6689 - 划分名称: 测试集, 字节数: 486076659.954, 样本量: 6726 下载总大小: 4707509912, 数据集总大小: 5590026631.847 配置项: - 配置名称: 默认配置, 数据文件: - 训练集: data/train-* - 验证集: data/validation-* - 测试集: data/test-* 任务类别: - 自动语音识别 - 会话类 语言: 西班牙语 样本量范围: 1万至10万 友好显示名称: COSER-ASR 子集 # 简介 “COSER-ASR 子集”是从《西班牙乡村口语有声语料库》(Corpus Oral y Sonoro del Español Rural,简称COSER;Fernández-Ordóñez 2005年至今)中提取的专业子集,该语料库意为“有声乡村西班牙语语料库”。本数据集专为微调自动语音识别系统Whisper而打造,已从原始COSER语料库中自动截取3至30秒的音频与文本片段。这些片段涵盖了多样且精炼的乡村西班牙语口语样本,非常适合用于训练与优化语音识别模型。为兼顾可处理性与计算效率,数据集单条样本的Token(Token)数量上限设为1024,在覆盖广度与计算成本间取得了平衡。 # 内容与人口统计聚焦 原始COSER语料库包含218份半结构化访谈转录文本,受访对象主要为西班牙乡村地区的老年、低教育水平群体。每份访谈时长平均约54分钟,蕴含丰富的方言变体与语言细节,为研究传统西班牙语方言提供了宝贵资料。 # 转写方案 本数据集提供多层级转写方案,以满足不同语言学研究与计算应用的需求: ## 原始转写(sentence) 该字段完整保留音频片段的原始口语转录,尽可能还原原始语音内容与完整的原始转写文本。 ## 语音近似转写(sentence_fono) 该版本调整了转写内容,以反映方言发音的语音学特征,对于理解乡村西班牙语方言的语音细节至关重要。 ## 无标记语音转写(sentence_fono_sin_marcas) 此转写移除了通常以方括号标注的话语标记(如笑声、应答声等),仅保留口语词汇,提供更简洁的文本版本。 ## 标准正写法对应(sentence_orto) 该层级提供语音转写内容对应的标准正写法文本,搭建了方言口语与标准西班牙语正写法之间的桥梁。 ## 无标记标准正写法转写(sentence_orto_sin_marcas) 与无标记语音转写类似,该版本移除了所有话语标记,提供纯净的标准正写法文本,特别适用于需要干净文本数据的应用场景。 # 局限性 本数据集存在以下局限:COSER语料库中的时间区间未经过系统对齐,因此音频与文本数据之间可能无法实现完全的一一对应。 # 补充信息与资源 若需了解更多COSER语料库的研究方法与完整转写内容,可访问http://coser.lllf.uam.es/ 与 http://coser.lllf.uam.es/transcripcion.php。这些资源将深入展示COSER项目如何全面捕捉乡村西班牙语的语言多样性。 # 参考文献 Fernández-Ordóñez, I.(编辑). (2005年至今). 《西班牙乡村口语有声语料库》. 2022年4月15日检索自 http://www.corpusrural.es/
数据集概述
数据集信息
-
特征列表:
audio: 音频数据filename: 文件名turno_id: 标识符turno_time: 时间sentence: 原始转录sentence_fono: 音系近似转录sentence_fono_sin_marcas: 无话语标记的音系转录sentence_orto: 正字法对应转录sentence_orto_sin_marcas: 无话语标记的正字法转录Provincia: 省份Enclave: 飞地Fecha: 日期Duración: 持续时间Informantes: 发音人
-
数据分割:
train: 训练集,包含 53971 个样本,大小为 4600923777.433 字节validation: 验证集,包含 6689 个样本,大小为 503026194.46 字节test: 测试集,包含 6726 个样本,大小为 486076659.954 字节
-
数据集大小:
- 下载大小: 4707509912 字节
- 数据集大小: 5590026631.847 字节
-
配置:
default:- 训练集路径:
data/train-* - 验证集路径:
data/validation-* - 测试集路径:
data/test-*
- 训练集路径:
-
任务类别:
- 自动语音识别
- 对话
-
语言:
- 西班牙语
-
数据集名称:
- COSER-ASR Subset
-
数据集规模:
- 10K<n<100K
内容和人口统计重点
原始COSER数据集包括218份半结构化访谈,主要对象是来自西班牙农村的老年、受教育程度较低的个体。这些访谈平均时长约54分钟,富含方言变体和语言细节,为传统西班牙方言提供了宝贵的见解。
转录方法
- 原始转录 (sentence): 直接转录音频片段,尽可能保留原始语音和完整原始转录。
- 音系近似转录 (sentence_fono): 转录修改以反映方言发音的音系特征,对于理解农村西班牙方言的音韵细节至关重要。
- 无话语标记的音系转录 (sentence_fono_sin_marcas): 去除话语标记(如笑声、同意等),提供仅关注口语的更清洁版本。
- 正字法对应转录 (sentence_orto): 提供音系转录的标准正字法等价物,弥合方言语音和标准西班牙语正字法之间的差距。
- 无话语标记的正字法转录 (sentence_orto_sin_marcas): 类似于无标记的音系转录,提供无话语标记的标准正字法文本,特别适用于需要清洁文本数据的应用。
局限性
该模型的局限性包括COSER语料库中的时间间隔未系统对齐,这意味着音频和文本数据之间可能不存在完美的一一对应关系。




