ViF-CoT-4K
收藏资源简介:
ViF-CoT-4K是由清华大学团队构建的首个大规模AI生成视频人工标注数据集,包含4000条精细标注的样本。该数据集涵盖十余种前沿视频生成模型(如Sora-2、Wan2.2等)的生成内容,通过分层标注体系对视频中的物理异常和低阶伪造痕迹进行系统分类。数据来源包括真实视频(采集自Panda-70M等公开数据集)和AI生成视频,经过严格的语义对齐和人工标注流程构建。该数据集专为可解释AI视频检测任务设计,通过提供时空定位的伪造证据,支持模型在内容安全、数字取证等领域的应用。
ViF-CoT-4K is the first large-scale manually annotated dataset of AI-generated videos developed by the Tsinghua University research team, consisting of 4000 finely annotated samples. This dataset covers the generated content from over ten cutting-edge video generation models such as Sora-2, Wan2.2 and others, and systematically classifies physical anomalies and low-level forgery traces in videos via a hierarchical annotation framework. Its data sources include real videos (collected from public datasets including Panda-70M) and AI-generated videos, and it is constructed through rigorous semantic alignment and manual annotation pipelines. This dataset is specifically tailored for explainable AI video detection tasks, and supports model applications in fields such as content security and digital forensics by providing spatially and temporally localized forgery evidence.
Skyra数据集概述
数据集名称
ViF-CoT-4K
数据集简介
ViF-CoT-4K是一个专为可解释的AI生成视频检测而构建的大规模数据集。它旨在解决现有数据集中缺乏详细伪影标注的问题,为模型的监督微调(SFT)提供支持。
核心特征
- 规模:包含约4,000个视频。
- 视频来源:包含来自Sora-2、Wan2.1、Kling等模型的高质量AI生成视频样本。
- 标注内容:提供细粒度的人工标注,包括伪影类型、文本解释、时间戳和边界框。
- 数据设计:生成的视频与真实视频在语义上对齐,构成“真实-伪造”配对,以防止模型进行捷径学习。
构建目的
该数据集用于训练Skyra模型,这是一个专注于“基于伪影的推理”的多模态大语言模型(MLLM),旨在通过识别人类可感知的视觉伪影,为AI生成视频检测提供可解释的依据。
获取方式
训练数据ViF-CoT-4K可从以下地址下载:https://huggingface.co/datasets/JoeLeelyf/ViF-CoT-4K
许可证
ViF-CoT-4K数据集在CC BY 4.0许可证下发布。使用者必须遵守其源数据集(Kinetics-400, Panda-70M, HD-VILA-100M)的条款。

- 1Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning清华大学 · 2025年



