遇见数据集

Prompt2SceneGallery: A Visual Gallery of Indoor Scenes Generated from Structured Prompt Templates

收藏
Zenodo2025-07-29 更新2026-05-26 收录
官方服务:

资源简介:

Prompt2SceneGallery dataset showcases one of the utilities of the Prompt2SceneBench Dataset. The dataset consists of 5163 indoor scene images, the images were generated using Stable Diffusion XL (SDXL), and the prompts were randomly sampled from the Prompt2SceneBench Dataset. The images generated by SDXL have dimensions of 1024x1024. There are 4 types of prompts, therefore we have 4 types of images:Type A: 1 ObjectType B: 2 ObjectsType C: 3 ObjectsType D: 4 Objects Each row in the CSV corresponds to a single prompt instance and includes the following fields: type: Prompt category — one of A, B, C, or D, based on number of objects and complexity. object1, object2, object3, object4: Objects involved in the scene (some may be None/NaN/Null depending on type). surface: The surface where the objects are placed (e.g., desk surface, bench). scene: The indoor environment (e.g., living room, study room). prompt: The final structured natural language prompt. filename: Name of the image file. To get a detailed understanding, please go through the description of the Prompt2SceneBench Dataset. Direct Use Cases for Prompt2SceneGallery: Prompt–Image Alignment Evaluation (Imperfect Realism) Analyze how closely generated images match structured prompts in terms of object presence, co-location, and scene context — even when generation is imperfect. Failure Case Analysis for Text-to-Image Models Study the failure modes of models like SDXL in spatial reasoning, compositionality, or object fidelity. Visual Grounding Benchmarking (With Noise) Use imperfect generations to stress-test grounding models on scene understanding under visually noisy conditions. Evaluation Dataset for Captioning Models Evaluate how well captioning models (e.g., BLIP, LLaVA) can describe structured scenes — including their limitations in hallucinated or partially wrong outputs. Robustness and Semantic Drift Studies Explore how generation quality affects semantic drift between prompt and image, especially for structured spatial prompts. Synthetic Scene Prototyping for Research Serve as a starting point for prototyping indoor spatial benchmarks without needing human annotation.

提供机构:
Zenodo
创建时间:
2025-07-29
二维码
社区交流群
二维码
科研交流群
商业服务