VACaps: A Domain-Specific Image-Caption Dataset for Video Analytics
收藏官方服务:
资源简介:
An image–caption dataset comprising 30k images, 210k captions, and 304k bounding box annotations, designed to support Video Analytics Research and Applications. It contains real-world images depicting activities and objects relevant to surveillance, public safety, and abnormal behavior detection. The dataset aims to facilitate research in: Vision-Language Modeling Object Detection Models Multimodal Reasoning Development of Video Analytic Applications for (e.g., Text-based Image Retrieval using multimodals) Dataset Size: 30,000 images 210,000 captions (seven captions per image) 304,043 bounding box annotations (multiple bounding boxes per image)
提供机构:
Zenodo创建时间:
2026-06-13



