pictorlabs/pathbench
收藏资源简介:
PathBench是一个视觉组织病理学多项选择题(MCQ)基准测试,用于评估医学视觉语言模型在组织病理学图像解释上的性能。数据集包含30个问题,每个问题配有一张来自Wikimedia Commons(CC BY-SA许可)的最大800px的JPEG图像,涵盖12个病理学亚专业(如胃肠道病理学、乳腺病理学、皮肤病理学、神经病理学、泌尿生殖病理学、血液病理学等),并按难度分为简单(9题)、中等(16题)和困难(5题)。每个问题采用6选项多项选择格式,图像为H&E或IHC染色。数据集以JSONL格式提供,包含问题ID、图像路径、来源、器官、亚专业、染色方法、难度、问题描述、选项、答案和解释等字段。该基准测试支持通过API端点进行评估,并提供了基线模型(如Claude Opus 4.6和MedGemma)的性能结果。数据集主要用于医学AI模型的研究和评估,但具有小规模、单图像、英语仅限等局限性。
PathBench is a visual histopathology multiple-choice question (MCQ) benchmark developed to evaluate the performance of medical vision-language models in histopathological image interpretation. The dataset contains 30 questions, each paired with a JPEG image up to 800px in size sourced from Wikimedia Commons under the CC BY-SA license. It covers 12 pathological subspecialties, such as gastrointestinal pathology, breast pathology, dermatopathology, neuropathology, genitourinary pathology, hematopathology, and others. The questions are classified into three difficulty levels: simple (9 questions), moderate (16 questions) and hard (5 questions). Each question adopts a 6-option multiple-choice format, and the images are stained with hematoxylin and eosin (H&E) or immunohistochemistry (IHC). The dataset is provided in JSON Lines (JSONL) format, including fields such as question ID, image path, source, organ, subspecialty, staining method, difficulty level, question description, options, correct answer and explanation. This benchmark supports evaluation via API endpoints, and provides performance results of baseline models like Claude Opus 4.6 and MedGemma. The dataset is mainly used for research and evaluation of medical AI models, but has limitations such as small scale, single-image per question and English-only restriction.





