JAMMEval
收藏资源简介:
JAMMEval是由日本多所顶尖研究机构联合构建的日语视觉语言模型评估基准,通过对7个现有日语VQA数据集的系统性优化形成。该数据集包含1,592个精炼样本,覆盖OCR、日本文化知识、多图像推理等七大领域,采用两轮人工标注流程修正了原始数据中模糊问题、错误答案等缺陷。其创新性在于通过重标注而非简单过滤来保证数据规模,显著提升了评估信度,适用于检验模型对日语多模态任务的真实理解能力。
JAMMEval is a Japanese visual-language model evaluation benchmark jointly constructed by multiple leading Japanese research institutions, formed through systematic optimization of seven existing Japanese Visual Question Answering (VQA) datasets. This dataset contains 1,592 refined samples covering seven domains including OCR, Japanese cultural knowledge, multi-image reasoning and more, and adopts a two-round manual annotation process to correct defects such as ambiguous questions and incorrect answers in the original data. Its core innovation lies in ensuring data scale through re-annotation rather than simple filtering, which significantly improves evaluation reliability and is applicable to testing the real understanding capabilities of models for Japanese multimodal tasks.
JAMMEval 数据集概述
数据集名称
JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation
核心贡献者
- Issa Sugiura (Kyoto University, NII LLMC)
- Koki Maeda (Institute of Science Tokyo, NII LLMC)
- Shuhei Kurita (NII, NII LLMC)
- Yusuke Oda (NII LLMC)
- Daisuke Kawahara (Waseda University, NII LLMC)
- Naoaki Okazaki (Institute of Science Tokyo, NII LLMC)
资源链接
- 论文地址:https://speed1313.github.io/JAMMEval (arXiv)
- 代码地址:https://speed1313.github.io/JAMMEval (Code)
- 博客地址:https://speed1313.github.io/JAMMEval (Blog)
研究背景与目标
现有的日语视觉问答评估数据集存在局限性,包括问题表述模糊、偶尔的标注不准确,以及某些问题无需视觉基础即可回答的情况,这降低了评估的可靠性。为解决这些问题,构建了JAMMEval。
数据集描述
JAMMEval是一个经过重新标注的评估数据集集合,源自七个广泛使用的日语基准测试。所有实例都经过两轮人工审查和重新标注,以产生一个精炼的基准集合。
主要特点与影响
- 通过重新标注,提高了区分模型性能的分辨率。
- 在精炼后,所有模型的准确率都有所提高,且多次运行间的方差减小。
- 移除了有问题的实例,并将模糊的问题替换为客观上可回答的问题,从而获得更稳定和可靠的评估分数。
评估结果
在JAMMEval的七个任务上评估了现有模型。Gemini 3 Pro(启用了推理功能)总体得分最高,而所有其他模型均在未启用推理功能的情况下进行评估。

- 1JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation京都大学; 日本国立情报学研究所·LLMC; 日本国立情报学研究所; 早稻田大学; 东京科学研究所 · 2026年



