mlvbench-review/MLV-Bench
收藏资源简介:
MLV-Bench是一个用于野外长上下文医学视频理解的基准测试。它包含来自8个公共来源的340个公开完整医学视频,总计759解码小时,以及1,253个经过验证的多选题。该数据集旨在评估多模态模型在长上下文医学视频理解、稀疏证据检索和多跳推理方面的性能,不适用于临床诊断、患者管理或部署。数据集采用JSONL格式,每个视频记录包含嵌套的QA项。此外,还提供了一个代表性的样本供审阅者检查数据质量。
MLV-Bench is a benchmark for long-context medical video understanding in the wild. It contains 340 public full-procedure medical videos from 8 public sources, totaling 759 decoded hours, and 1,253 verified multiple-choice questions. The dataset is intended for research evaluation of multimodal models on long-context medical video understanding, sparse evidence retrieval, and multi-hop reasoning. It is not intended for clinical diagnosis, patient management, or deployment. The dataset is structured in JSONL format, with each video record containing nested QA items. A representative sample is also provided for reviewers to inspect data quality.



