VISTA
收藏资源简介:
VISTA是由犹他大学的一组研究人员创建的视觉和文本注意力数据集,旨在探索和增强视觉语言模型(VLMs)的可解释性。该数据集通过手动标注的方式,将图像区域与相应的文本段落进行对齐,记录了志愿者的眼球运动和语音描述,最终形成了508个高质量的图像-文本对齐数据。VISTA数据集的创建过程严格遵循伦理标准,确保了数据的隐私和匿名性。该数据集主要用于研究视觉语言模型中的注意力机制,旨在提高模型的透明度和可解释性,特别是在图像与文本的关联分析中。
VISTA is a visual and textual attention dataset developed by a team of researchers at the University of Utah, designed to explore and enhance the interpretability of Vision-Language Models (VLMs). This dataset aligns image regions with corresponding text passages via manual annotation, records eye movements and verbal descriptions from volunteers, and ultimately yields 508 high-quality image-text alignment pairs. The development process of the VISTA dataset strictly adheres to ethical standards, ensuring data privacy and anonymity. This dataset is primarily used to study attention mechanisms in vision-language models, aiming to improve model transparency and interpretability, particularly in image-text association analysis.

- 1VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models犹他大学 · 2024年



