遇见数据集

Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026

收藏
Zenodo2026-05-28 更新2026-05-26 收录
官方服务:

资源简介:

The Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026 is a structured dataset preserving English-language transcript records from the Samuel & Audrey travel video archive. The dataset is part of the Samuel & Audrey Media Network research archive and provides machine-readable transcript data for long-running independent travel media published across YouTube and related web platforms. It is designed to support research into travel media, creator-economy history, tourism communication, video metadata, transcript analysis, multilingual and multimodal media archives, retrieval-based AI research, and long-form travel storytelling. The corpus includes transcript text and associated metadata from Samuel & Audrey travel videos published between 2012 and 2026. Depending on the source files included in this package, records may contain video titles, URLs or identifiers, publication dates, transcript text, cue-level or segment-level transcript fields, channel context, language information, and source/provenance metadata. This dataset should be interpreted as a historical media transcript corpus rather than a complete representation of every video ever published by Samuel & Audrey. Some transcripts may be generated, corrected, segmented, normalized, or otherwise prepared for research and retrieval workflows. Users should consult the included README, data dictionary, schema, manifest, and citation files for field definitions, package structure, limitations, licensing, and citation guidance. The current canonical working version is maintained through the Samuel & Audrey Media Network dataset stack, with related mirrors on Hugging Face, GitHub, Zenodo, Kaggle, DagsHub, and other research-data platforms where applicable.

提供机构:
Zenodo
创建时间:
2026-02-17
二维码
社区交流群
二维码
科研交流群
商业服务