COPACAYO: Corpus del Parlamento de Canarias en YouTube
收藏资源简介:
COPACAYO: Corpus del Parlamento de Canarias en YouTube Project Context COPACAYO (Corpus del Parlamento de Canarias en YouTube) is a comprehensive corpus of Spanish-language YouTube videos from Canarian parliamentary and governmental institutions containing timestamped transcriptions. This dataset was developed as part of project A09 "On the interplay between register and socio-geographic variation in Canarian Spanish" within the Collaborative Research Centre 1412 "REGISTER" (Register: Language Users' Knowledge of Situational-Functional Variation), led by Prof. Dr. Miriam Bouzouita at Humboldt-Universität zu Berlin. The corpus spans over 14 years of video content from major Canarian governmental institutions, parliamentary sessions, and official channels, providing a valuable resource for linguistic research, natural language processing, and computational linguistics studies focusing on Spanish language evolution and regional institutional discourse patterns. **Project Information**: [A09 - On the interplay between register and socio-geographic variation in Canarian Spanish](https://sfb1412.hu-berlin.de/projects/a09/) Dataset Characteristics Basic Information - **Language**: Spanish (Canarian Spanish) - **Content Type**: YouTube video transcriptions with timestamps from Canarian governmental institutions - **Total Videos**: 5,744 - **Total Files**: 1 Parquet file - **Valid Dates**: 5,744 videos (100% coverage) - **Temporal Coverage**: December 12, 2009 - May 31, 2024 - **Duration**: 5,284 days (14.5 years) - **Total Duration**: 77.5 days, 11 hours, 18 minutes, 47 seconds - **Total Tokens**: 16,929,289 tokens Data Structure The Parquet file contains 5,744 videos with the following metadata: - **Title**: Video title - **Author**: YouTube channel name (governmental institution) - **Length**: Video duration in seconds - **Views**: View count - **Video URL**: Direct YouTube link - **Video ID**: Unique YouTube identifier - **Publish Date**: Publication date - **Tags**: Video tags - **Description**: Video description - **Transcription**: Timestamped transcription in format `[HH:MM:SS] text content` - **Transcription_no_timestamp**: Clean transcription without timestamps - **Transcription_punct**: Transcription with punctuation Institutional Coverage The corpus includes content from major Canarian governmental institutions: - **Cabildo de Lanzarote y La Graciosa**: 46.4% of content (2,664 videos) - **Cabildo de Tenerife**: 30.4% of content (1,748 videos) - **Cabildo de La Gomera**: 11.2% of content (646 videos) - **Cabildo de Gran Canaria**: 5.6% of content (322 videos) - **Parlamento de Canarias**: 1.8% of content (103 videos) - **Cabildo Fuerteventura**: 1.6% of content (92 videos) - **Cabildo de Tenerife - Live**: 1.1% of content (64 videos) - **Isla de El Hierro**: 1.0% of content (56 videos) - **Cabildo de La Palma en directo**: 0.9% of content (49 videos) Content Characteristics - **Average Video Duration**: 19.4 minutes (1,165 seconds) - **Total Views**: 2,538,884 views - **Average Views per Video**: 442 views - **Transcription Coverage**: 100% of videos have transcriptions - **Content Types**: Parliamentary sessions, governmental meetings, official announcements, institutional communications License This dataset was created from YouTube subtitles for linguistic and computational analysis. The data were extracted under the text and data mining exception (§44b UrhG, Directive 2019/790/EU) applicable in Germany, for non-commercial scientific research purposes. Redistribution of the original texts is restricted due to copyright. Therefore, the dataset is archived under closed access. Derived annotations, metadata, and scripts can be made available upon request for academic collaboration. Contact For questions, issues, or collaboration requests, please contact j.bonilla@hu-berlin.de. Version History - **v1.0** (2025): Initial release with 5,744 videos spanning 2009-2024



