VISFACTOR
收藏资源简介:
VISFACTOR是一个基于FRCT的视觉认知基准测试,由香港中文大学、香港中文大学(深圳)和腾讯AI实验室共同开发。该数据集包含15个视觉和空间推理测试,旨在评估大型多模态语言模型在空间推理、感知速度和模式识别等核心视觉认知任务上的能力。测试内容涵盖了从寻找隐藏图形、复制模式到空间扫描和视觉辨别等多种任务,每个任务都针对特定的视觉认知能力。VISFACTOR为研究人员提供了一个自动化的评估框架,以及人类表现基准,以推动在大型多模态语言模型视觉认知能力领域的研究。
VISFACTOR is a FRCT-based visual cognition benchmark jointly developed by The Chinese University of Hong Kong, The Chinese University of Hong Kong, Shenzhen, and Tencent AI Lab. This dataset comprises 15 visual and spatial reasoning tests, designed to evaluate the capabilities of large multimodal language models on core visual cognitive tasks including spatial reasoning, perceptual speed, and pattern recognition. The test contents cover a variety of tasks such as finding hidden figures, reproducing patterns, spatial scanning, and visual discrimination, with each task targeting a specific visual cognitive ability. VISFACTOR provides researchers with an automated evaluation framework and human performance benchmarks to advance research in the field of visual cognitive capabilities of large multimodal language models.




