AEGIS
收藏资源简介:
AEGIS(即**A**ssessing **E**diting, **G**eneration, **I**nterpretation-**U**nderstanding for **S**uper-intelligence)是一个全面的多任务基准测试,旨在评估统一多模态模型(UMMs)在多样化任务中应用世界知识的能力。该数据集包含1,050个具有挑战性的手动标注问题,涵盖21个主题(包括STEM、人文、日常生活等)和6种推理类型。AEGIS通过视觉理解、生成、编辑和交错生成任务,具体评估UMMs在世界知识范围内的表现,并提出了确定性基于检查表的评估(DCE)协议,以增强评估的可靠性。实验结果表明,大多数UMMs在世界知识方面存在严重缺陷,且性能随着推理复杂度的增加而显著下降。此外,简单的插件推理模块可以部分缓解这些缺陷,为未来研究指明了方向。
AEGIS (short for **A**ssessing **E**diting, **G**eneration, **I**nterpretation-**U**nderstanding for **S**uper-intelligence) is a comprehensive multi-task benchmark designed to evaluate the ability of unified multimodal models (UMMs) to apply world knowledge across diverse tasks. This dataset comprises 1,050 challenging manually annotated questions, covering 21 topics (including STEM, humanities, daily life, etc.) and six types of reasoning. AEGIS specifically assesses the performance of UMMs within the scope of world knowledge through visual understanding, generation, editing and interleaved generation tasks, and proposes a Deterministic Checklist-based Evaluation (DCE) protocol to enhance the reliability of the assessment. Experimental results demonstrate that most UMMs suffer from severe deficiencies in world knowledge, and their performance declines significantly as the complexity of reasoning increases. Furthermore, simple plug-in reasoning modules can partially alleviate these deficiencies, providing clear directions for future research.
AEGIS数据集概述
数据集基本信息
- 数据集名称: AEGIS
- 许可证: MIT
- 任务类别: 图像-文本到文本、视觉问答、问答、文本到图像
- 支持语言: 英语、中文
- 数据规模: 1K<n<10K
数据集简介
AEGIS(Assessing Editing, Generation, Interpretation-Understanding for Super-intelligence)是一个用于评估统一多模态模型(UMMs)世界知识应用能力的综合性多任务基准。该基准旨在解决现有基准在评估上的局限性,通过涵盖视觉理解、生成、编辑和交错生成等多种任务,对模型进行全面的诊断性评估。
核心特点
- 多任务评估: 同时评估视觉理解、生成、编辑和交错生成能力。
- 广泛的知识覆盖: 包含1,050个具有挑战性的手动标注问题,涵盖21个主题(包括STEM、人文学科、日常生活等)和6种推理类型。
- 确定性评估协议: 提出了基于检查表的确定性评估(DCE)协议,使用原子化的“是/否”判断替代模糊的基于提示的评分,以提高评估的可靠性。
- 深入诊断: 揭示了当前先进统一多模态模型存在的严重世界知识缺陷,以及推理复杂性对性能的影响。
数据构成
- 领域覆盖: 涵盖STEM、人文学科和日常生活三大领域,包含21个多样化主题。
- 问题分布: 每个主题包含15个用于视觉理解、生成和编辑的提示,以及5个用于衡量复杂生成能力的视觉交错生成问题。
- 推理类型: 大多数提示中融入了六种不同的推理类型,要求模型具备内在的推理能力来完成请求。
相关资源
- GitHub仓库: https://github.com/DongSky/AEGIS
- 博客文章: https://m1saka.moe/aegis/
- 论文地址: https://arxiv.org/abs/2601.00561
引用信息
如果本工作对您的研究有所帮助,请考虑引用: bibtex @misc{aegis, title={AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models}, author={Jintao Lin, Bowen Dong, Weikang Shi, Chenyang Lei, Suiyun Zhang, Rui Liu, Xihui Liu}, year={2026}, eprint={2601.00561}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2601.00561}, }




