SK-VG
收藏资源简介:
SK-VG数据集由香港中文大学(深圳)创建,旨在通过场景知识引导视觉定位任务,提升模型对图像和文本的推理能力。该数据集包含约40,000个参考表达式和8,000个场景故事,源自4,000张图像,每张图像包含两个场景故事和五个相关表达式。数据集的创建过程涉及图像选择和手动注释,确保数据的高质量和多样性。SK-VG数据集的应用领域广泛,特别是在视觉问答和视觉导航等任务中,有助于评估和提升机器的视觉与语言理解能力。
The SK-VG dataset was created by The Chinese University of Hong Kong, Shenzhen, aiming to enhance models' reasoning abilities for images and texts by leveraging scene knowledge to guide visual grounding tasks. This dataset contains approximately 40,000 referring expressions and 8,000 scene stories, sourced from 4,000 images, with each image paired with two scene stories and five corresponding referring expressions. The dataset construction process involves image selection and manual annotation to ensure high data quality and diversity. SK-VG has a wide range of application scenarios, especially in tasks such as visual question answering (VQA) and visual navigation, where it helps evaluate and improve machines' visual and language understanding capabilities.




