SceneEdit3D-15K
收藏资源简介:
SceneEdit3D-15K是由华东师范大学、上海人工智能实验室等机构联合构建的大规模三维场景编辑数据集,旨在为标准化评估提供高质量的配对监督数据。该数据集包含15,000个精心配对的场景编辑样本,每个样本均提供了完整的RGB视频、深度图、相机参数、编辑掩码以及语言指令等多模态标注信息,数据来源于基于Blender渲染的可组合室内场景。其创建过程通过视觉语言模型提出规范的编辑任务,经可行性检查和渲染管线生成配对数据,确保了三维结构的一致性。该数据集主要应用于三维场景编辑算法的训练与评估,致力于解决现有方法缺乏配对三维标注、难以量化评估编辑保真度与几何一致性的核心瓶颈问题。
SceneEdit3D-15K is a large-scale 3D scene editing dataset jointly constructed by East China Normal University, Shanghai AI Laboratory and other institutions, aiming to provide high-quality paired supervised data for standardized evaluation. This dataset contains 15,000 meticulously paired scene editing samples, each of which provides multi-modal annotation information including complete RGB videos, depth maps, camera parameters, editing masks and language instructions. The data is sourced from compositional indoor scenes rendered based on Blender. Its creation process proposes standardized editing tasks via visual-language models, generates paired data through feasibility checks and rendering pipelines, ensuring the consistency of 3D structures. This dataset is mainly applied to the training and evaluation of 3D scene editing algorithms, and is committed to resolving the core bottleneck problem that existing methods lack paired 3D annotations and struggle to quantitatively evaluate editing fidelity and geometric consistency.
数据集名称
SceneEdit3D,包含两个主要部分:
- SceneEdit3D-15K:15K 个配对的编辑样本,覆盖场景级物体移除、移动、外观编辑和混合操作。
- SceneEdit3D-Bench:一个精选的 100 样本基准测试集,用于评估编辑保真度、背景保持和3D结构质量。
数据集内容
- 数据来源:从可组合的3D室内场景出发,使用视觉语言模型(VLM)提出规范编辑,在 Blender 中执行编辑,并渲染配对的原/目标视频。
- 标注信息:包含 RGB、深度、相机参数、掩码、编辑参考帧和指令监督,以及渲染器提供的几何、相机、掩码和辅助监督,支持RGB-几何编辑和评估。
- 样本格式:每对样本包含原视频与编辑后视频、半透明掩码覆盖、编辑提示示例,以及原/编辑后的3D点云预览。
数据集用途
用于训练和标准化评估前馈式3D场景编辑方法,特别是为论文提出的 JointEdit3D 框架提供训练数据和评测基准。
数据规模与特点
- 15K 配对编辑样本,覆盖5种操作类型:删除、添加、移动、外观编辑和混合操作。
- 提供渲染器输出的3D标注,包括几何、相机、掩码等,支持 RGB-几何联合编辑和三维结构质量的评估。
- 基准测试包含 100 个样本,用于衡量编辑区域质量、背景保持能力和3D结构完整性。




