HOIVG-Bench
收藏资源简介:
HOIVG-Bench是由香港中文大学、字节跳动等机构联合构建的综合性评估基准,专注于人-物交互视频生成(HOIVG)任务的多模态条件验证。该数据集通过整合文本提示、参考图像、音频和姿态序列等异构数据,填补了该领域标准化评估工具的空白。其设计采用严格的五元组数据配对(条件输入与目标视频),支持对生成视频的视觉质量、跨模态对齐精度等核心指标进行系统化测评。作为OmniShow框架的配套基准,它为解决电商展示、短视频制作等实际应用中的可控视频生成难题提供了量化研究基础。
HOIVG-Bench is a comprehensive evaluation benchmark jointly constructed by institutions including The Chinese University of Hong Kong and ByteDance, focusing on multimodal conditional validation for human-object interaction video generation (HOIVG) tasks. This dataset integrates heterogeneous data such as text prompts, reference images, audio and pose sequences, filling the gap of standardized evaluation tools in this field. Its design adopts a strict quintuple data pairing scheme (conditional input and target video), enabling systematic evaluation of core metrics including the visual quality of generated videos and cross-modal alignment accuracy. As a supporting benchmark for the OmniShow framework, it provides a quantitative research foundation for addressing the challenges of controllable video generation in practical applications such as e-commerce display and short-video production.
OmniShow 数据集概述
基本描述
- 数据集名称: OmniShow
- 核心任务: 人-物交互视频生成
- 技术特点: 一个端到端的框架,用于统一文本、参考图像、音频和姿态条件以合成高质量的人-物交互视频。
主要功能与模式
- 参考图像到视频生成: 通过注入参考图像,实现高保真外观和自然交互。
- 参考图像+音频到视频生成: 在音频输入下,保持参考身份并与音频更可靠地对齐运动。
- 参考图像+姿态到视频生成: 给定参考图像和姿态,更好地遵循运动轨迹,同时保持物体交互的真实性。
- 参考图像+音频+姿态到视频生成: 独特地支持联合文本+参考+音频+姿态输入,实现精确条件对齐的稳定生成。
技术优势
- 逼真的运动质量: 具有丰富且连贯动态的平滑运动。
- 稳健的物理合理性: 更稳定的接触、抓握和更少的穿透。
- 原生长镜头生成: 生成更长的连续镜头,最长可达10秒。
- 富有表现力的虚拟形象动画: 从人物图像和音频输入生成生动的说话和唱歌。
- 稳定的身份保持: 在不同场景中保持高度一致的角色外观。
对比基准
- 在参考图像到视频生成任务中,与 HunyuanCustom、HuMo-17B、VACE 和 Phantom-14B 进行了比较。
- 在参考图像+音频到视频生成任务中,与 HunyuanCustom 和 HuMo-17B 进行了比较。
- 在参考图像+姿态到视频生成任务中,与 AnchorCrafter 和 VACE 进行了比较。
相关资源
- 论文链接: https://correr-zhou.github.io/OmniShow
- GitHub链接: https://correr-zhou.github.io/OmniShow
- 评估基准: HOIVG-Bench
- 更多相关工作: HiFi-Inpaint, IdentityStory, SceneDecorator, MagicTailor

- 1OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation香港中文大学; 字节跳动; 莫纳什大学; 香港大学 · 2026年



