ThumbnailTruth
收藏资源简介:
ThumbnailTruth数据集是一个多模态数据集,包含来自八个国家(四个发展中国家和四个发达国家)的2,843个YouTube视频,其中1,359个视频的缩略图具有误导性。这些误导性缩略图视频的总观看量超过76亿次,为研究跨文化背景下误导性缩略图问题提供了独特的视角。数据集包括视频缩略图、视频到文本描述和字幕转录等信息,用于全面分析内容并标记误导性缩略图。该数据集旨在帮助研究和开发用于检测YouTube等平台上的误导性缩略图的模型,从而提高内容质量和用户信任度。
ThumbnailTruth Dataset is a multimodal dataset containing 2,843 YouTube videos from eight countries (four developing and four developed), among which 1,359 videos feature misleading thumbnails. The total views of these videos with misleading thumbnails exceed 7.6 billion, providing a unique perspective for studying the issue of misleading thumbnails across cultural contexts. The dataset encompasses video thumbnails, video-to-text descriptions, subtitle transcriptions and other relevant information, enabling comprehensive content analysis and labeling of misleading thumbnails. This dataset is designed to support research and development of models for detecting misleading thumbnails on platforms such as YouTube, thereby enhancing content quality and user trust.

- 1ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings拉合尔管理科学大学计算机科学系 · 2025年



