ATBench
收藏资源简介:
ATBench是由卡尔斯鲁厄理工学院创建的一个专为视觉语言模型在辅助技术领域评估而设计的多模态基准数据集。该数据集包含五个与视觉语言任务相关的核心任务,包括全景分割、深度估计、光学字符识别、图像描述生成和视觉问答。数据集通过与视觉障碍者的用户研究相结合,确保了任务的相关性和实用性。创建过程中,研究团队通过问卷调查和用户反馈,筛选出对视觉障碍者最为重要的任务。ATBench旨在评估和提升视觉语言模型在辅助技术中的应用,特别是为视觉障碍者提供更全面的场景理解和辅助功能。
ATBench is a multimodal benchmark dataset developed by Karlsruhe Institute of Technology, specifically designed for evaluating vision-language models in the domain of assistive technology. It encompasses five core tasks related to vision-language paradigms, including panoptic segmentation, depth estimation, optical character recognition (OCR), image captioning, and visual question answering (VQA). To ensure the relevance and practicality of the tasks, the dataset integrates user studies conducted with visually impaired individuals. During its development, the research team screened out the most critical tasks for visually impaired populations through questionnaires and user feedback. The core objective of ATBench is to evaluate and advance the application of vision-language models in assistive technology, particularly to provide more comprehensive scene understanding and assistive functionalities for visually impaired people.




