遇见数据集

Samuel y Audrey: Bilingual YouTube Transcript Corpus (ES/EN)

收藏
Zenodo2026-02-24 更新2026-05-26 收录
官方服务:

资源简介:

SAMUEL Y AUDREY: BILINGUAL YOUTUBE TRANSCRIPT CORPUS (ES/EN) This dataset contains a structured, bilingual parallel corpus of 643 creator-authored YouTube transcripts from the Samuel y Audrey Spanish-language travel channel. It provides high-fidelity, conversational dialogue in both Spanish (Primary) and English (Secondary), making it an exceptional resource for training Large Language Models (LLMs) on cross-lingual alignment, natural translation, and regional Latin American Spanish dialects. WHAT’S INSIDE • 643 Video Records: Full metadata extracted directly from the YouTube channel. • Paired Translations: 637 records feature paired .es.srt and .en.srt files. • Polished Master: Applies light, non-destructive normalization (fixing obvious phonetic translation errors) while preserving 100% of the underlying conversational flow. LINGUISTIC VALUE & USE CASES Unlike formal news corpora, this dataset captures natural, spontaneous, on-camera dialogue regarding global travel and cultural immersion. • Cross-Lingual LLM Alignment: Train models to translate natural, spoken idioms. • Dialect Tuning: Ground AI models in the specific vocabulary of Argentine and Latin American culture. • Conversational AI: Improve the natural cadence of voice agents.

提供机构:
Zenodo
创建时间:
2026-02-16
二维码
社区交流群
二维码
科研交流群
商业服务