HumaniBench
收藏资源简介:
HumaniBench是一个包含32K真实世界图像-问题对的综合基准,通过可扩展的GPT-4o辅助流程进行注释,并由领域专家进行彻底验证。HumaniBench通过七个不同的任务探索了七个HCAI原则——公平性、道德、理解、推理、语言包容性、同理心和鲁棒性,这些任务包括开放式和封闭式视觉问答(VQA)、多语言QA、视觉定位、同理性标题和鲁棒性测试。对15个最先进的LMMs(开源和闭源)的基准测试表明,专有模型通常领先;然而,在鲁棒性和视觉定位方面仍存在一些差距,而一些开源模型在平衡准确性与遵守人类对齐原则(如道德和包容性)方面存在困难。HumaniBench是第一个围绕HCAI原则构建的基准。它提供了一个严格的测试平台,用于诊断对齐差距,并引导LMMs朝着既准确又负责任的社会行为发展。为了促进透明度和支持未来的研究,我们发布了数据集、注释提示和评估代码。
HumaniBench is a comprehensive benchmark comprising 32K real-world image-question pairs, annotated via a scalable GPT-4o-assisted pipeline and thoroughly validated by domain experts. HumaniBench explores seven Human-Centered AI (HCAI) principles—fairness, ethics, understanding, reasoning, linguistic inclusivity, empathy, and robustness—across seven distinct tasks, including open-ended and closed-ended visual question answering (VQA), multilingual QA, visual grounding, empathetic captioning, and robustness testing. Benchmarking 15 state-of-the-art large multimodal models (LMMs, both open-source and closed-source) shows that proprietary models generally lead; however, notable gaps remain in robustness and visual grounding, while some open-source models struggle to balance accuracy with adherence to human-aligned principles such as ethics and inclusivity. HumaniBench is the first benchmark built around HCAI principles. It provides a rigorous testbed for diagnosing alignment gaps and guiding large language models (LLMs) toward both accurate and responsible social behavior. To promote transparency and support future research, we have publicly released the dataset, annotation prompts, and evaluation code.

- 1HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation多伦多,加拿大的Vector Institute和奥兰多,美国的中央佛罗里达大学 · 2025年



