Corpus Mediático de Canarias en YouTube
收藏资源简介:
# COMECAYO: Corpus mediático de Canarias en YouTube (Corpus of Canary Islands Media on YouTube) ## Overview Comecayo (Corpus mediático de Canarias en YouTube) is a comprehensive corpus of Spanish-language YouTube videos from Canarian media outlets containing timestamped transcriptions. This dataset was developed as part of project A09 "On the interplay between register and socio-geographic variation in Canarian Spanish" within the Collaborative Research Centre 1412 "REGISTER" (Register: Language Users' Knowledge of Situational-Functional Variation). The corpus spans over 15 years of video content from major Canarian television channels, radio stations, and digital media platforms, providing a valuable resource for linguistic research, natural language processing, and computational linguistics studies focusing on Spanish language evolution and regional media discourse patterns. **Project Information**: [A09 - On the interplay between register and socio-geographic variation in Canarian Spanish](https://sfb1412.hu-berlin.de/projects/a09/) ## Dataset Characteristics ### Basic Information - **Language**: Spanish (Canarian Spanish) - **Content Type**: YouTube video transcriptions with timestamps from Canarian media outlets - **Total Videos**: 38,947 - **Total Files**: 7,790 Parquet files - **Valid Dates**: 37,766 videos (97% coverage) - **Temporal Coverage**: December 14, 2008 - August 18, 2024 - **Duration**: 5,726 days (15.7 years) - **Total Duration**: 694 days, 2 hours, 52 minutes, 41 seconds - **Total Tokens**: 228,102,778 tokens ### Data Structure Each Parquet file contains approximately 5 videos with the following metadata: - **Title**: Video title - **Author**: YouTube channel name - **Length**: Video duration in seconds - **Views**: View count - **Video URL**: Direct YouTube link - **Video ID**: Unique YouTube identifier - **Publish Date**: Publication date - **Tags**: Video tags - **Description**: Video description - **Transcription**: Timestamped transcription in format `[HH:MM:SS] text content` ### Media Outlets Coverage The corpus includes content from major Canarian media outlets: - **InformativosTvc**: 63.4% of content (primary news source) - **TelevisionCanaria**: 14.0% of content (regional television) - **Lancelot Digital**: 7.7% of content (digital media) - **Mírame TV**: 6.1% of content (television channel) - **Canarias Noticias**: 3.8% of content (news outlet) - **CANAL4TENERIFE**: 2.4% of content (local television) - **Canarias Radio**: 1.0% of content (radio station) - **TV La Palma**: 0.8% of content (local television) - **CANAL 11 LA PALMA OFICIAL**: 0.6% of content (local television) - **Noticanarias**: 0.1% of content (news outlet) ## Technical Specifications ### File Format - **Primary Format**: Apache Parquet - **Encoding**: UTF-8 - **Compression**: Snappy compression - **Schema**: Structured with 10 columns per record ### Data Quality - **Completeness**: 97% of videos have valid publication dates - **Consistency**: Standardized timestamp format `[HH:MM:SS]` - **Validation**: Automated quality checks for date parsing and text normalization ## Data Access ### File Organization ``` corpus/ ├── corpus_parquet/ # Main data files (7,790 files) │ ├── corpus_data_part_1.parquet │ ├── corpus_data_part_2.parquet │ └── ... ├── corpus_txt/ # Text-only versions └── statistics.xlsx # Corpus statistics ``` ## Citation If you use this corpus in your research, please cite: ```bibtex @dataset{comecayo2024, title={COMECAYO: Corpus mediático de Canarias en YouTube (Corpus of Canary Islands Media on YouTube)}, author={Bonilla, Johnatan E.}, year={2024}, publisher={Zenodo}, doi={[DOI will be assigned]}, url={[Zenodo URL]} } ``` ## License This dataset was created from YouTube subtitles for linguistic and computational analysis. The data were extracted under the text and data mining exception (§44b UrhG, Directive 2019/790/EU) applicable in Germany, for non-commercial scientific research purposes. Redistribution of the original texts is restricted due to copyright. Therefore, the dataset is archived under closed access. Derived annotations, metadata, and scripts can be made available upon request for academic collaboration. ## Contact For questions, issues, or collaboration requests, please contact [contact information]. ## Acknowledgments The Corpus mediático de Canarias en YouTube was compiled by Johnatan E. Bonilla as part of project A09 "On the interplay between register and socio-geographic variation in Canarian Spanish" of the Collaborative Research Centre 1412 "REGISTER" (Register: Language Users' Knowledge of Situational-Functional Variation), led by Prof. Dr. Miriam Bouzouita at Humboldt-Universität zu Berlin. We acknowledge the Canarian media outlets whose content is included in this dataset: InformativosTvc, TelevisionCanaria, Lancelot Digital, Mírame TV, Canarias Noticias, CANAL4TENERIFE, Canarias Radio, TV La Palma, CANAL 11 LA PALMA OFICIAL, and Noticanarias. We also acknowledge the open-source community for the tools used in data processing. ## Version History - **v1.0** (2024): Initial release with 38,947 videos spanning 2008-2024 --- **Keywords**: COMECAYO, Corpus mediático de Canarias, Canarian Spanish, Canarian media, YouTube, transcriptions, timestamps, linguistic corpus, regional Spanish, media discourse, temporal analysis, natural language processing, computational linguistics, Canary Islands



