UCA (UCF-Crime Annotation)
收藏资源简介:
UCA数据集是由北京工业大学等机构合作创建的,专注于监控视频与语言理解的第一个多模态数据集。该数据集包含23,542个句子级别的描述,平均长度为20个单词,标注视频总时长达到110.7小时。UCA数据集通过精细的事件内容和时间标注,支持多种多模态理解任务,如视频字幕生成、密集视频字幕和多模态异常检测,旨在解决监控视频内容自动理解的挑战,提升现有调查措施在监控应用中的效能。
The UCA dataset, co-created by institutions including Beijing University of Technology, is the first multimodal dataset dedicated to surveillance video and language understanding. It contains 23,542 sentence-level descriptions, with an average length of 20 words, and the total duration of the annotated videos reaches 110.7 hours. With fine-grained event content and temporal annotations, the UCA dataset supports a variety of multimodal understanding tasks, such as video captioning, dense video captioning, and multimodal anomaly detection. It aims to address the challenges in automatic understanding of surveillance video content and improve the efficiency of existing investigative measures in surveillance applications.

- 1Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges北京工业大学 · 2023年



