OphthalWeChat
收藏资源简介:
OphthalWeChat是一个大规模的双语(中英)视觉问答(VQA)基准,专为眼科设计。该资源旨在支持在现实眼科场景中对视觉语言模型(VLM)进行开发和严格评估,以实现自动化诊断、医学教育和远程医疗等应用。数据集包括3,469张眼科图像和30,120个问答对,涵盖9个眼科亚专科、29种成像模态和68种模态组合。OphthalWeChat是第一个双语VQA基准,具有现实世界背景,并包含每个患者的多次检查,反映了真实的临床决策场景,并使VLMs的定量评估成为可能,支持开发准确、专业和值得信赖的AI系统,以用于眼科护理。
OphthalWeChat is a large-scale bilingual (Chinese-English) visual question answering (VQA) benchmark designed specifically for ophthalmology. This resource aims to support the development and rigorous evaluation of vision-language models (VLMs) in real-world ophthalmological scenarios, enabling applications such as automated diagnosis, medical education, and telemedicine. The dataset includes 3,469 ophthalmic images and 30,120 question-answer pairs, covering 9 ophthalmic subspecialties, 29 imaging modalities, and 68 modality combinations. OphthalWeChat is the first bilingual VQA benchmark with real-world context, which contains multiple examinations per patient, reflecting authentic clinical decision-making scenarios and enabling quantitative evaluation of VLMs, thus supporting the development of accurate, specialized, and trustworthy AI systems for ophthalmic care.

- 1Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat香港理工大学 · 2025年



