遇见数据集

IVUL-KAUST/MedCTA

收藏
Hugging Face2026-05-24 更新2026-06-14 收录
官方服务:

资源简介:

MedCTA是一个用于评估临床工具代理的基准数据集。每个数据样本包含一张临床图像、一个临床用户查询、一个参考工具使用轨迹和一个最终的真实答案。该数据集旨在评估代理在医疗环境中的多模态能力,包括理解临床图像和图表、选择适当的工具、检索或提取证据、在需要时进行计算、跨工具调用整合观察结果,以及回答基于临床的问题。数据集包含107个样本,涉及5种工具(如计算器、OCR、图像描述、谷歌搜索和区域属性描述),平均每个样本进行3.2次工具调用和8.38轮对话。它支持视觉问答、问答、图像到文本和文本生成等任务,适用于医疗、临床人工智能、工具使用、代理、多模态和基准测试等领域的研究。

MedCTA is a benchmark dataset for evaluating clinical tool agents. Each data sample consists of a clinical image, a clinical user query, a reference tool usage trajectory, and a final ground-truth answer. This dataset is designed to assess the multimodal capabilities of agents within clinical environments, including comprehending clinical images and charts, selecting appropriate tools, retrieving or extracting evidence, conducting computations when required, integrating observations across multiple tool calls, and responding to clinical-oriented questions. It contains 107 samples covering 5 categories of tools, such as calculators, OCR, image captioning, Google Search, and regional attribute description, with an average of 3.2 tool calls and 8.38 dialogue turns per sample. This dataset supports tasks including visual question answering, question answering, image-to-text generation, and text generation, and is applicable to research in fields such as healthcare, clinical artificial intelligence, tool usage, AI agents, multimodality, and benchmark testing.

提供机构:
IVUL-KAUST
二维码
社区交流群
二维码
科研交流群
商业服务