LENS
收藏资源简介:
LENS是一个多层级基准测试,专门设计用于评估多模态大型语言模型(MLLMs)在感知、理解和推理三个层次上的表现。它包含8个任务和12个真实场景,拥有3.4K当代照片和40K领域特定问答对,全部由人工标注并经领域专家审核。图像多样且新颖,约53%的图像来自2025年,80%以上的图像来自2024年9月之后。数据集还包括300多个细粒度视觉类别和65种不同的问题风格,旨在评估功能视觉智能和推理能力。
LENS is a multi-level benchmark specially designed to evaluate the performance of Multimodal Large Language Models (MLLMs) across three core dimensions: perception, understanding, and reasoning. It comprises 8 tasks and 12 real-world scenarios, encompassing 3.4K contemporary photographs and 40K domain-specific question-answer pairs, all manually annotated and reviewed by domain experts. The images are diverse and novel, with approximately 53% sourced from 2025, and over 80% collected after September 2024. The dataset also includes over 300 fine-grained visual categories and 65 distinct question styles, aiming to assess functional visual intelligence and reasoning capabilities.
LENS 多模态大型语言模型评估基准数据集
📌 数据集概述
- 名称: LENS (Large-scale Evaluation Benchmark for Multimodal LLMs)
- 类型: 多模态评估基准
- 设计目标: 通过三级评估体系(感知、理解、推理)评估MLLMs,涵盖8项任务和12个现实场景
🌟 核心特征
-
数据规模
- 3.4K张当代照片
- 40K个领域特定问答对(人工标注+专家审核)
-
数据时效性
- 53%图像来自2025年
- 80%以上图像采集于2024年9月之后
- 覆盖12个现实场景
-
标注粒度
- 300+细粒度视觉类别
- 65+种问题类型
🎯 评估任务
-
感知层任务
- 描述性物体计数
- 物体检测(300+细粒度类别)
- 物体存在性判定
-
理解层任务
- 关系提取(如"持有"、"相邻"等)
- 视觉定位(自然语言→图像区域)
- 区域OCR(指定区域内文字识别)
-
推理层任务
- 空间关系理解
- 场景知识推理
📊 基准性能对比
| 模型(Qwen2.5-VL) | 参数量 | 感知得分 | 理解得分 | 推理得分 |
|---|---|---|---|---|
| 3B版本 | 3B | 0.5876 | 0.6652 | 0.6075 |
| 7B版本 | 7B | 0.5835 | 0.7158 | 0.7061 |
| 32B版本 | 32B | 0.6225 | 0.7457 | 0.5166 |
| 72B版本 | 72B | 0.5975 | 0.7598 | 0.7095 |
⚠️ 使用声明
- 所有人脸图像均已进行隐私保护处理
- 数据集发布状态: 待论文接收后公开(当前未发布)




