VideoPath-Instruct
收藏资源简介:
VideoPath-Instruct是一个包含4278个视频和诊断相关思维链对的数据集,来源于YouTube上的教育性组织病理学视频。该数据集旨在模仿病理学家的自然诊断过程,包含单张切片图像、自动提取的关键帧视频片段和手动分割的视频病理图像三种不同的图像场景。虽然高质量数据对增强诊断推理至关重要,但其创建过程耗时且数据量有限。为了克服这一挑战,我们从现有的单图像指令数据集中迁移知识,以训练弱注释的关键帧提取视频,然后手动分割视频进行微调。VideoPath-LLaVA在病理学视频分析中建立了新的基准,并为未来支持临床决策的AI系统提供了有前景的基础。
VideoPath-Instruct is a dataset containing 4278 video-diagnosis related chain-of-thought pairs, sourced from educational histopathology videos on YouTube. This dataset aims to mimic the natural diagnostic process of pathologists, and includes three distinct image scenarios: single-slide images, automatically extracted keyframe video clips, and manually segmented video pathological images. While high-quality data is crucial for enhancing diagnostic reasoning, its creation is time-consuming and the volume of available data is limited. To address this challenge, we transferred knowledge from existing single-image instruction datasets to train weakly annotated keyframe-extracted videos, followed by manual segmentation of videos for fine-tuning. VideoPath-LLaVA establishes a new benchmark in pathological video analysis and provides a promising foundation for future AI systems that support clinical decision-making.
VideoPath-LLaVA 数据集概述
数据集基本信息
- 名称: VideoPath-LLaVA
- 领域: 病理学诊断推理
- 技术方向: 视频指令调优
数据集用途
- 用于病理学诊断推理的视频指令调优研究。
数据集特点
- 提供病理学诊断推理的视频数据。
- 包含指令调优的示例和结果。
相关资源
- 论文: VideoPath-LLaVA: Pathology Diagnostic Reasoning Through Video Instruction Tuning
- 项目页面: VideoPath-LLaVA Project Page
引用信息
bibtex @misc{vuong2025videopathllavapathologydiagnosticreasoning, title={VideoPath-LLaVA: Pathology Diagnostic Reasoning Through Video Instruction Tuning}, author={Trinh T. L. Vuong and Jin Tae Kwak}, year={2025}, eprint={2505.04192}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2505.04192}, }

- 1VideoPath-LLaVA: Pathology Diagnostic Reasoning Through Video Instruction Tuning韩国大学 · 2025年



