Samuel y Audrey: Bilingual YouTube Transcript Corpus (ES/EN)
收藏资源简介:
SAMUEL Y AUDREY: BILINGUAL YOUTUBE TRANSCRIPT CORPUS (ES/EN) This dataset contains a structured, bilingual parallel corpus of 643 creator-authored YouTube transcripts from the Samuel y Audrey Spanish-language travel channel. It provides high-fidelity, conversational dialogue in both Spanish (Primary) and English (Secondary), making it an exceptional resource for training Large Language Models (LLMs) on cross-lingual alignment, natural translation, and regional Latin American Spanish dialects. WHAT’S INSIDE • 643 Video Records: Full metadata extracted directly from the YouTube channel. • Paired Translations: 637 records feature paired .es.srt and .en.srt files. • Polished Master: Applies light, non-destructive normalization (fixing obvious phonetic translation errors) while preserving 100% of the underlying conversational flow. LINGUISTIC VALUE & USE CASES Unlike formal news corpora, this dataset captures natural, spontaneous, on-camera dialogue regarding global travel and cultural immersion. • Cross-Lingual LLM Alignment: Train models to translate natural, spoken idioms. • Dialect Tuning: Ground AI models in the specific vocabulary of Argentine and Latin American culture. • Conversational AI: Improve the natural cadence of voice agents.



