COACAN: Corpus del español académico de Canarias
收藏资源简介:
COACAN: CORPUS OF ACADEMIC CANARIAN SPANISH PROJECT CONTEXT COACAN (Corpus of Academic Canarian Spanish) is a linguistic corpus that documents and analyzes varieties of Spanish spoken in the Canary Islands, specifically within the university academic context. This dataset was developed as part of project A09 "On the interplay between register and socio-geographic variation in Canarian Spanish" of the Collaborative Research Centre 1412 "REGISTER" (Register: Language Users' Knowledge of Situational-Functional Variation), led by Prof. Dr. Miriam Bouzouita at Humboldt-Universität zu Berlin. This corpus was developed at the University of La Laguna with students from three specific degree programs in October-November 2024: • Physical Activity and Sports Sciences (CAFYD)• Hispanic Linguistics and Literature• Education (Pedagogy) The corpus focuses on spontaneous speech from Canarian university students, capturing both dialectal features of Canarian Spanish and academic/colloquial registers used in higher education settings, as well as their relationship with insular territory. DATA COLLECTION METHODOLOGIES The COACAN corpus employs two complementary data collection methodologies: 1. HABLACANARIABOT METHODOLOGY (Manually revised transcriptions) Data collection was conducted using HablaCanariaBot (https://web.telegram.org/k/#@HablaCanariaBot) , a Telegram bot specifically designed for linguistic research. This automated system enabled: • Collection of spontaneous speech samples in natural, familiar environments • Transcriptions that were subsequently manually revised • Detailed sociolinguistic metadata collection The bot presented questions designed to elicit specific Canarian Spanish linguistic phenomena, including locative adverbs, existential "haber" usage, nominal possessives, and other characteristic dialectal features. 2. SOCIAL CARTOGRAPHY AND PRESENTATIONS METHODOLOGY (Automatic transcriptions) This methodology captures academic and territorial speech through: • Social cartography work: Students work independently creating maps about questions related to the island, their relationship with territory and perceptual dialectology • Formal group presentations: Structured academic expositions • Focus on territory-identity-language relationships CORPUS STRUCTURE COACAN comprises four subcorpora organized according to the two collection methodologies: HABLACANARIABOT METHODOLOGY (Manually revised transcriptions): 1. INDIVIDUAL CORPUS Format: Individual responses to directed questions Participants: 79 unique users Transcriptions: 635 recordings Content: Short monologues and direct responses 2. PAIRS CORPUS Format: Conversations between participant pairs Participants: 79 pairs (158 people) Transcriptions: 835 recordings Content: Spontaneous dialogues and guided conversations SOCIAL CARTOGRAPHY METHODOLOGY (Automatic transcriptions): 3. GROUPS CORPUS (SOCIAL CARTOGRAPHY) Format: Collaborative social cartography work Participants: 17 working groups Transcriptions: 14 work sessions Content: Group conversations during map creation Focus: Territory-identity-perceptual dialectology relationships 4. PRESENTATIONS CORPUS (FORMAL PRESENTATIONS) Format: Group academic presentations Participants: 17 presentation groups Transcriptions: 14 formal presentations Content: Structured academic discourse Focus: Formal register and academic competence DATA FORMAT AND STRUCTURE PARTICIPANT METADATA:• Demographics: age, gender, origin• Academic information: university, major, year of study• Geographic data: birthplace, upbringing, residence TRANSCRIPTION DATA:• Unique question identifier• Question text• Target linguistic phenomenon• Original audio file (.ogg)• Synchronized transcription (SRT format)• Audio duration and recording timestamp TECHNICAL FORMAT:• Structure: Hierarchical JSON• Encoding: UTF8• Audio: OGG Vorbis format• Transcriptions: SRT format with timestamps CORPUS STATISTICS • Total transcriptions: 1,498 recordings• HablaCanariaBot methodology: 1,470 transcriptions (manually revised)• Social cartography methodology: 28 transcriptions (automatic with Whisper)• Transcribed words: Over 50,000 tokens• Degree programs represented: CAFYD, Hispanic Linguistics and Literature, Education• Geographic coverage: Primarily Tenerife, with representation from other islands• Age range: 18-25 years (average: ~22 years) ETHICAL CONSIDERATIONS • Informed consent from all participants• GDPR compliance• Academic research use only• Controlled access for authorized researchers CONTACT INFORMATION For inquiries about the COACAN corpus, additional data access, or research collaborations, contact: Johnatan E. BonillaHumboldt-Universität zu Berlinj.bonilla@hu-berlin.de Document generated: 2025Corpus version: 1.0Last updated: October 2025



