遇见数据集

Cross-Lingual YouTube Metadata with AI Topic Labels (English and Portuguese)

收藏
Zenodo2026-02-27 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains metadata and topic annotations for AI-related YouTube videos collected in English (USA) and Portuguese (Brazil). The data was curated through a multi-stage pipeline including transcript retrieval, large language model-based screening, human validation, and supervised topic classification. Two parquet files are provided: videos_usa.parquet videos_brazil.parquet Each file includes the following fields: VideoId: Unique YouTube video identifier Title: Video title PublishedAt: Publication date ChannelId: Unique channel identifier Duration: Video duration topics_list: List of topic indexes assigned to the video topic: Dominant topic index topics_list_name: List of topic labels topic_name: Dominant topic label video_type: Video format category (e.g., short-form or medium-length) The dataset does not include raw transcripts due to copyright and privacy restrictions. It is intended for academic research on AI-related content, cross-lingual topic modeling, and temporal analysis of online discourse

提供机构:
Zenodo
创建时间:
2026-02-27
二维码
社区交流群
二维码
科研交流群
商业服务