FunBench
收藏资源简介:
FunBench是由中国人民大学AIMClab创建的一个视觉问题回答(VQA)基准,旨在全面评估多模态大型语言模型(MLLM)的视网膜阅读技能。该数据集包含16,348个视网膜图像和91,810个视觉问题,涵盖了从低级别的模态感知、解剖感知到高级别的病变分析和疾病诊断四个层次的任务。数据来源于多个公开数据集,包括彩色眼底摄影(CFP)、光学相干断层扫描(OCT)和超广角眼底摄影(UWF)等。FunBench的设计考虑了两个基本问题:问什么和怎么问,以实现对MLLM视网膜阅读技能的全面评估。该数据集应用于评估MLLM在眼科图像分析领域的性能,解决视网膜图像的解读问题。
FunBench is a visual question answering (VQA) benchmark created by the AIMClab of Renmin University of China, aiming to comprehensively evaluate the retinal reading skills of multimodal large language models (MLLMs). The dataset includes 16,348 retinal images and 91,810 visual questions, covering four hierarchical tasks ranging from low-level modality perception and anatomical perception to high-level lesion analysis and disease diagnosis. Its data is sourced from multiple public datasets, including color fundus photography (CFP), optical coherence tomography (OCT), ultra-widefield fundus photography (UWF), and other related modalities. The design of FunBench takes into account two core questions: what to ask and how to ask, enabling a comprehensive assessment of the retinal reading abilities of MLLMs. This dataset is utilized to evaluate the performance of MLLMs in the field of ophthalmic image analysis, addressing the challenge of retinal image interpretation.

- 1FunBench: Benchmarking Fundus Reading Skills of MLLMs中国人民大学 · 2025年



