ViTT (Video Timeline Tags)
收藏资源简介:
ViTT 数据集由人工制作的 8,169 个视频的片段级注释组成。其中,5840 个视频被注释一次,其余视频被注释两次或更多。共发布了 12461 组注解。数据集中的视频来自 Youtube-8M 数据集。 注释具有以下格式: { "id": "FmTp", “注释”:[ { “时间戳”:260, “标签”:“开幕” }, { “时间戳”:16000, "tag": "展示技巧" }, { “时间戳”:23990, "tag": "显示脚部定位" }, { “时间戳”:55530, "tag": "演示跨界" }, { “时间戳”:114100, “标签”:“关闭” } ] }
The ViTT dataset consists of segment-level annotations for 8,169 manually curated videos. Of these, 5,840 videos are annotated once, while the remaining videos are annotated two or more times. A total of 12,461 annotation sets have been released. The videos in the dataset are sourced from the Youtube-8M dataset. The annotations follow the format: { "id": "FmTp", "annotations": [ { "timestamp": 260, "label": "opening" }, { "timestamp": 16000, "label": "demonstration technique" }, { "timestamp": 23990, "label": "show foot positioning" }, { "timestamp": 55530, "label": "demonstrate crossover" }, { "timestamp": 114100, "label": "closing" } ] }




