COMECAYO: Corpus Mediático de Canarias en YouTube
收藏资源简介:
COMECAYO: Corpus mediático de Canarias en YouTube (Corpus of Canary Islands Media on YouTube) Project context Comecayo (Corpus mediático de Canarias en YouTube) is a comprehensive corpus of Spanish-language YouTube videos from Canarian media outlets containing timestamped transcriptions. This dataset was developed as part of project A09 "On the interplay between register and socio-geographic variation in Canarian Spanish" within the Collaborative Research Centre 1412 "REGISTER" (Register: Language Users' Knowledge of Situational-Functional Variation), led by Prof. Dr. Miriam Bouzouita at Humboldt-Universität zu Berlin.. The corpus spans over 15 years of video content from major Canarian television channels, radio stations, and digital media platforms, providing a valuable resource for linguistic research, natural language processing, and computational linguistics studies focusing on Spanish language evolution and regional media discourse patterns. **Project Information**: [A09 - On the interplay between register and socio-geographic variation in Canarian Spanish](https://sfb1412.hu-berlin.de/projects/a09/) Project A09 centers on the crossroads between register and socio-geographic variation in Canarian Spanish for three morpho-syntactic phenomena. The linguistic reality of the Canary Islands is complex as they belong politically to Spain and are influenced by the Peninsular Spanish standard variety, while they share many linguistic features with Latin American Spanish. The project aims to find evidence for a diaglossic situation, characterized by a dialect-to-standard continuum where intermediate varieties exist between the local dialects and the national standard variety, considering the interplay between register parameters (e.g., formality) and socio-geographic ones (e.g., education, age). Dataset Characteristics Basic Information - **Language**: Spanish (Canarian Spanish) - **Content Type**: YouTube video transcriptions with timestamps from Canarian media outlets - **Total Videos**: 38,947 - **Total Files**: 7,790 Parquet files - **Valid Dates**: 37,766 videos (97% coverage) - **Temporal Coverage**: December 14, 2008 - August 18, 2024 - **Duration**: 5,726 days (15.7 years) - **Total Duration**: 694 days, 2 hours, 52 minutes, 41 seconds - **Total Tokens**: 228,102,778 tokens Data Structure Each Parquet file contains approximately 5 videos with the following metadata: - **Title**: Video title - **Author**: YouTube channel name - **Length**: Video duration in seconds - **Views**: View count - **Video URL**: Direct YouTube link - **Video ID**: Unique YouTube identifier - **Publish Date**: Publication date - **Tags**: Video tags - **Description**: Video description - **Transcription**: Timestamped transcription in format `[HH:MM:SS] text content` Media Outlets Coverage The corpus includes content from major Canarian media outlets: - **InformativosTvc**: 63.4% of content (primary news source) - **TelevisionCanaria**: 14.0% of content (regional television) - **Lancelot Digital**: 7.7% of content (digital media) - **Mírame TV**: 6.1% of content (television channel) - **Canarias Noticias**: 3.8% of content (news outlet) - **CANAL4TENERIFE**: 2.4% of content (local television) - **Canarias Radio**: 1.0% of content (radio station) - **TV La Palma**: 0.8% of content (local television) - **CANAL 11 LA PALMA OFICIAL**: 0.6% of content (local television) - **Noticanarias**: 0.1% of content (news outlet) Data Access File Organization ``` corpus/ ├── corpus_parquet/ Main data files (7,790 files) │ ├── corpus_data_part_1.parquet │ ├── corpus_data_part_2.parquet │ └── ... License This dataset was created from YouTube subtitles for linguistic and computational analysis. The data were extracted under the text and data mining exception (§44b UrhG, Directive 2019/790/EU) applicable in Germany, for non-commercial scientific research purposes. Redistribution of the original texts is restricted due to copyright. Therefore, the dataset is archived under closed access. Derived annotations, metadata, and scripts can be made available upon request for academic collaboration. Contact For questions, issues, or collaboration requests, please contact j.bonilla@hu-berlin.de. Version History - **v1.0** (2025): Initial release with 38,947 videos spanning 2008-2024
# COMECAYO:加那利群岛YouTube媒体语料库(Corpus of Canary Islands Media on YouTube) ## 项目背景 COMECAYO(加那利群岛YouTube媒体语料库)是一套涵盖加那利群岛媒体机构发布的西班牙语YouTube视频的综合性语料库,包含带时间戳的转录文本。本数据集是德国柏林洪堡大学Miriam Bouzouita教授主导的合作研究中心1412「REGISTER」(意为「语域:语言使用者对情境功能变体的认知」)下属A09项目「论加那利西班牙语语域与社会地理变异的相互作用」的研究成果。本语料库收录了15年来加那利群岛主要电视频道、广播电台及数字媒体平台的视频内容,可为聚焦西班牙语演化与区域媒体话语模式的语言学研究、自然语言处理(Natural Language Processing,NLP)及计算语言学研究提供宝贵资源。 **项目信息**:[A09 - 论加那利西班牙语语域与社会地理变异的相互作用](https://sfb1412.hu-berlin.de/projects/a09/) 项目A09围绕加那利西班牙语中的三种形态句法现象,聚焦语域与社会地理变异的交叉研究。加那利群岛在政治上隶属于西班牙,受半岛西班牙语标准变体影响,同时又与拉丁美洲西班牙语共享诸多语言特征,其语言生态颇为复杂。本项目旨在探究双言现象的实证依据,该现象以方言-标准语连续体为特征,即存在本地方言与国家标准变体之间的中间变体,同时需考量语域参数(如正式程度)与社会地理参数(如受教育程度、年龄)之间的相互作用。 ## 数据集特征 ### 基本信息 - **语言**:西班牙语(加那利西班牙语) - **内容类型**:加那利群岛媒体机构发布的带时间戳的YouTube视频转录文本 - **视频总数**:38,947条 - **文件总数**:7,790个Parquet文件 - **有效日期覆盖**:37,766条视频(覆盖率97%) - **时间跨度**:2008年12月14日至2024年8月18日 - **总时长跨度**:5,726天(合15.7年) - **总内容时长**:694天2小时52分41秒 - **总Token数**:228,102,778个Token ### 数据结构 每个Parquet文件约包含5条视频,附带以下元数据: - **标题**:视频标题 - **作者**:YouTube频道名称 - **时长**:视频时长(单位:秒) - **播放量**:观看次数 - **视频链接**:YouTube直接访问链接 - **视频ID**:YouTube唯一标识符 - **发布日期**:视频发布时间 - **标签**:视频标签 - **描述**:视频简介 - **转录文本**:带时间戳的转录内容,格式为`[HH:MM:SS] 文本内容` ### 媒体机构覆盖范围 本语料库收录了加那利群岛主要媒体机构的内容: - **InformativosTvc**:占总内容的63.4%(核心新闻来源) - **TelevisionCanaria**:占总内容的14.0%(区域电视频道) - **Lancelot Digital**:占总内容的7.7%(数字媒体) - **Mírame TV**:占总内容的6.1%(电视频道) - **Canarias Noticias**:占总内容的3.8%(新闻机构) - **CANAL4TENERIFE**:占总内容的2.4%(本地电视频道) - **Canarias Radio**:占总内容的1.0%(广播电台) - **TV La Palma**:占总内容的0.8%(本地电视频道) - **CANAL 11 LA PALMA OFICIAL**:占总内容的0.6%(本地电视频道) - **Noticanarias**:占总内容的0.1%(新闻机构) ## 数据获取 ### 文件组织 corpus/ ├── corpus_parquet/ 主数据文件(共7,790个) │ ├── corpus_data_part_1.parquet │ ├── corpus_data_part_2.parquet │ └── ... ### 许可证 本数据集源自YouTube字幕,用于语言学与计算分析。数据提取遵循德国《版权法》第44b条及欧盟2019/790号指令规定的文本与数据挖掘例外条款,仅用于非商业性科学研究。 由于版权限制,原始文本的再分发受到约束,因此本数据集采用受限访问模式存档。 衍生注释、元数据及脚本可应学术合作请求提供。 ## 联系方式 如有疑问、问题或合作意向,请联系j.bonilla@hu-berlin.de。 ## 版本历史 - **v1.0**(2025年):初始版本,收录38,947条视频,时间跨度为2008年至2024年。



