HOI-Edit
收藏资源简介:
HOI-Edit是由北京大学研究团队构建的专用于评估人-物交互图像编辑任务的多层次认知基准数据集。该数据集包含705个精心标注的高质量样本,涵盖丰富的交互类别,数据来源于HOI视频数据集的高分辨率帧,并经过严格的人工标注与指令设计。其创建过程涉及三个核心步骤:基于场景上下文的手动指令编写、针对交互主体与客体的系统化区域边界框标注,以及面向评估需求的多维度问题构建。该数据集主要应用于计算机视觉与人工智能领域,旨在解决现有图像编辑方法在动态人-物交互建模上的不足,通过分层认知评估(基础编辑、空间理解、因果与物理推理)推动模型在保持实体身份一致性的同时,精确合成符合逻辑与物理规律的动态交互过程。
HOI-Edit is a multi-level cognitive benchmark dataset dedicated to evaluating human-object interaction (HOI) image editing tasks, developed by a research team from Peking University. This dataset comprises 705 high-quality, meticulously annotated samples spanning diverse interaction categories, with its data sourced from high-resolution frames of HOI video datasets and subjected to rigorous manual annotation and instruction design. Its development involves three core steps: manual instruction writing based on scene context, systematic regional bounding box annotation for interaction subjects and objects, and multi-dimensional question construction tailored to evaluation requirements. Primarily applied in the fields of computer vision and artificial intelligence, this dataset aims to address the shortcomings of existing image editing methods in dynamic human-object interaction modeling. It advances model capabilities to accurately synthesize dynamic interaction processes that conform to logical and physical laws while maintaining entity identity consistency through hierarchical cognitive evaluations including basic editing, spatial understanding, causal and physical reasoning.
数据集:HOI-Edit
概述
HOI-Edit 是一个专注于人-物交互(Human-Object Interactions, HOI) 图像编辑的综合性基准数据集。它针对现有图像编辑方法在处理动态交互关系时的不足而构建,强调对动态交互有效性与人-物对保存的综合评估。
核心特征
- 三个渐进式认知层级:数据集按难度划分,用于测试模型在不同复杂度下的交互编辑能力。
- 自动化评估指标 HOI-Eval:首次通过视觉语言模型(VLM) 对包含定位人-物对的图像进行“思考后问答”,从而可靠地评估实例级别的交互有效性。
基准测试定位
- 挑战现有方法:指出当前图像编辑方法擅长静态属性编辑,但在复杂人-物交互编辑上表现不佳。
- 提出新框架 SCPE:基于数据集的评测,提出了自校正流程编辑(Self-Correcting Process Editing, SCPE) 框架。该框架利用图像到视频(I2V)模型的时序生成能力进行动态编辑,并通过迭代优化提示词(prompts)自我纠正错误,最终从生成视频中提取帧作为编辑结果。
数据集资源
- 论文下载:
assets/HOIEdit_ICML_2026_CR_submission.pdf - 概览图下载:
assets/fig1_v3.pdf - 示例概览图:
assets/fig1_v3.png
注意:当前数据集尚未发布(标记为“Coming Soon”)。





