遇见数据集

kepeng/MeasL-Bench-V1

收藏
Hugging Face2026-05-27 更新2026-05-31 收录
官方服务:

资源简介:

MeasL-Bench(全称:基于测量的视觉语言基准测试)是PRSIMVL项目发布的官方基准测试集,用于评估基于测量域输入的视觉语言模型。该数据集旨在测试一个核心问题:当RGB渲染丢失传感器证据时,基于测量域输入的模型是否能更可靠地恢复和推理。数据集强调低光照证据恢复、HDR和曝光敏感的基础、可见性和幻觉敏感查询,并包含标准RGB基础切片以确保覆盖范围。它包含2,183个示例,分为RAW测量路径和RGB对应版本,适用于评估RGB原生视觉语言模型和基于测量的方法(如Meas.-XYZ管道)。数据集按能力维度组织,包括颜色属性基础、数量基础、描述性场景基础等14个类别,并提供了评估协议和使用指南。

MeasL-Bench (short for Measurement-grounded Language-Vision Benchmark) is the official held-out benchmark released with PRSIMVL for measurement-grounded vision-language evaluation. It is designed to test a core question: when RGB rendering loses sensor evidence, can a model grounded on measurement-domain input recover and reason more reliably? The benchmark emphasizes low-illumination evidence recovery, HDR and exposure-sensitive grounding, visibility-sensitive and hallucination-sensitive queries, plus standard RGB-sufficient grounding slices for coverage. It contains 2,183 examples, with matched RAW and RGB versions, and is suitable for evaluating both RGB-native VLMs and measurement-grounded methods (e.g., Meas.-XYZ pipelines). It is organized by capability dimensions such as Chromatic Attribute Grounding, Numerosity Grounding, and others, and includes evaluation protocols and usage instructions.

提供机构:
kepeng
二维码
社区交流群
二维码
科研交流群
商业服务