ClevrTex: A Texture-Rich Benchmark for Unsupervised Multi-Object Segmentation
收藏资源简介:
There has been a recent surge in methods that aim to decompose and segment scenes into multiple objects in an unsupervised manner, i.e., unsupervised multi-object segmentation. Performing such a task is a long-standing goal of computer vision, offering to unlock object-level reasoning without requiring dense annotations to train segmentation models. Despite significant progress, current models are developed and trained on visually simple scenes depicting mono-colored objects on plain backgrounds. The natural world, however, is visually complex with confounding aspects such as diverse textures and complicated lighting effects. In this study, we present a new benchmark called ClevrTex, designed as the next challenge to compare, evaluate and analyze algorithms. ClevrTex features synthetic scenes with diverse shapes, textures and photo-mapped materials, created using physically based rendering techniques. ClevrTex has 50k examples depicting 3-10 objects arranged on a background, created using a catalog of 60 materials, and a further test set featuring 10k images created using 25 different materials. We benchmark a large set of recent unsupervised multi-object segmentation models on ClevrTex and find all state-of-the-art approaches fail to learn good representations in the textured setting, despite impressive performance on simpler data. We also create variants of the ClevrTex dataset, controlling for different aspects of scene complexity, and probe current approaches for individual shortcomings. Project webpage: https://www.robots.ox.ac.uk/~vgg/data/clevrtex/ This is the <strong>main dataset and OOD test set. </strong>Please see project page for links to the dataset variants.
近年来,旨在以无监督方式将场景分解并分割为多个物体的方法(即无监督多目标分割(unsupervised multi-object segmentation))迎来了爆发式增长。此类任务是计算机视觉领域长期以来的核心目标之一,有望实现无需密集标注即可训练分割模型的物体级推理。尽管已取得显著进展,但当前模型均基于视觉上较为简单的场景开发与训练,这类场景通常仅包含纯色物体与纯色背景。然而,真实世界的视觉场景极为复杂,存在多样纹理、复杂光照等诸多干扰因素。 本研究提出了一款名为ClevrTex的全新基准数据集,旨在作为后续挑战场景,用于对比、评估与分析各类算法。ClevrTex采用基于物理的渲染(physically based rendering)技术生成,包含多样化形状、纹理与照片级映射材质的合成场景。该数据集包含5万张训练样本,每张样本均在背景上排布3至10个物体,基于60种材质库生成;此外还附带1万张测试集样本,采用25种不同材质生成。 我们在ClevrTex上对多款近期无监督多目标分割模型开展了基准测试,结果发现,尽管这些方法在简单数据上表现亮眼,但在带纹理的场景中均无法学习到优质的特征表示。 我们还构建了ClevrTex数据集的多个变体,通过控制场景复杂度的不同维度,以探究当前各类方法存在的个体缺陷。项目网页:https://www.robots.ox.ac.uk/~vgg/data/clevrtex/ 此为**核心数据集与分布外(OOD)测试集**。如需获取数据集变体的链接,请参阅项目页面。



