遇见数据集

OpenRaiser/CoW-Bench

收藏
Hugging Face2026-05-26 更新2026-06-14 收录
官方服务:

资源简介:

CoW-Bench是一个全面的基准测试数据集,用于评估视频和图像生成模型对世界组合(CoW)的理解能力,重点关注多模态内容生成中的空间关系、对象交互和时间动态。该数据集挑战模型理解对象之间的复杂空间关系、生成准确表示对象遮挡和分层的内容、保持视频生成的时间一致性,并处理多对象交互和转换。它评估了视频生成模型(如Sora、Kling)和图像生成模型(如GPT-4V、DALL-E)在生成内容时是否遵守物理和空间约束。数据集包含1,435个样本,每个样本包括类别(如Space-2d-layed)、输入文本(场景的文本描述)、输入图像描述、输入图像(PIL图像对象)和输出视频期望(预期视频输出的描述)。输入模态为文本加图像,输出模态为视频或图像,评估方法基于文本与预期输出的比较。数据集旨在促进多模态生成模型的性能评估和优化。

CoW-Bench is a comprehensive benchmark dataset for evaluating video and image generation models understanding of Composition of World (CoW), focusing on spatial relationships, object interactions, and temporal dynamics in multi-modal content generation. It challenges models to understand complex spatial relationships between objects, generate content that accurately represents object occlusions and layering, maintain temporal consistency in video generation, and handle multi-object interactions and transformations. The benchmark evaluates both video generation models (e.g., Sora, Kling) and image generation models (e.g., GPT-4V, DALL-E) on their ability to generate content that respects physical and spatial constraints. The dataset contains 1,435 samples, each including a category (e.g., Space-2d-layed), input text (text description of the scenario), input image description, input image (PIL Image object), and output video expectation (description of expected video output). The input modality is text plus image, the output modality is video or image, and the evaluation method is based on text-based comparison with expected outputs. The dataset aims to facilitate performance evaluation and optimization of multi-modal generation models.

提供机构:
OpenRaiser
二维码
社区交流群
二维码
科研交流群
商业服务