leggedrobotics/funthor-dataset
收藏资源简介:
FunTHOR是一个用于功能3D场景理解的合成数据集,基于AI2-THOR模拟器构建。它提供了12个室内场景(如厨房、客厅、卧室、浴室)的部分级地面真实几何和密集的、基于规则的功能关系标注(例如“刀切苹果”、“把手拉开门”、“炉灶旋钮打开/关闭燃烧器”)。数据集包含每个场景的位姿RGB-D序列,总共有621个地面真实节点(对象+功能部件),其中92个是功能部件,以及164个功能关系边。每个场景有60个位姿RGB-D帧(分辨率1200×680),从可达视点随机采样。此外,数据集还提供对象和部件中心点云、对象-部件层次结构,以及每个场景的可见子集(仅包含从采样RGB-D帧中可观察的节点和边)。标注通过可检查的规则自动生成,覆盖部件-对象关系和对象-对象关系,旨在支持概率性、开放词汇功能3D场景图的构建。
FunTHOR is a synthetic dataset for functional 3D scene understanding, built on top of the AI2-THOR simulator. It provides part-level ground-truth geometry and dense, rule-based functional-relation annotations (e.g., knife slices apple, handle pulls to open door, stove knob turns on/off burner) for 12 indoor scenes, together with posed RGB-D sequences for each scene. The dataset includes 621 ground-truth nodes total (objects + functional parts), with 92 being functional parts, and 164 functional-relation edges. Each scene has 60 posed RGB-D frames (1200×680) randomly sampled from reachable viewpoints. It also offers object- and part-centric point clouds, an object-part hierarchy per scene, and a visible subset per scene that retains only nodes/edges observable from the sampled RGB-D frames. Annotations are generated automatically from inspectable rules, covering both part-object and object-object relations, and designed as a benchmark for constructing probabilistic, open-vocabulary functional 3D scene graphs.




