GENRES
收藏资源简介:
GENRES是一个旨在评估多模态大型语言模型(MLLMs)中性别偏见的新基准。它通过社会关系中的叙述来评估性别偏见,包括双角色配置文件和叙述生成任务,以捕捉丰富的人际动态,并支持跨多个维度的细致偏见评估。数据集包含1440个叙述引出对(NEPs),每个对都包括一个文本提示和相应的图像,描绘了一个男性和一个女性角色之间的社会互动。这些场景涵盖了不同的年龄、领域和关系动态,确保了全面评估。数据集的创建经历了四个阶段:叙述元素设计、NEP生成、响应收集和评估。评估方法整合了LLM和NLP工具,以评估角色配置文件和生成叙述中的偏见。GENRES旨在解决多模态生成系统中性别偏见的评估和缓解问题。
GENRES is a novel benchmark designed to evaluate gender bias in multimodal large language models (MLLMs). It assesses gender bias through narratives of social relationships, including dual-role profiles and narrative generation tasks, to capture rich interpersonal dynamics and enable fine-grained bias evaluation across multiple dimensions. The dataset comprises 1,440 narrative elicitation pairs (NEPs), each consisting of a textual prompt and a corresponding image depicting social interactions between a male and a female character. These scenarios cover diverse ages, domains, and relational dynamics to ensure comprehensive evaluation. The dataset construction involves four stages: narrative element design, NEP generation, response collection, and evaluation. The evaluation method integrates LLMs and NLP tools to assess bias in character profiles and generated narratives. GENRES aims to address the evaluation and mitigation of gender bias in multimodal generative systems.




