遇见数据集

Voxel51/KubriCount

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

KubriCount 是一个大规模合成基准数据集,用于多粒度视觉计数,由论文《Count Anything at Any Granularity》(Liu, Wu & Xie, SJTU 2026)提出。它将开放世界计数重新定义为跨五个明确语义粒度级别的提示跟随问题,并支持迄今为止发布的最全面注释的计数数据集。该数据集包含6736个样本,通过FiftyOne平台提供。每个场景是一个1024×1024的合成图像,通过四阶段自动管道生成:使用Kubric和Blender进行可控3D渲染、基于掩码的图像编辑以减少模拟到真实的差距,以及基于视觉语言模型的质量过滤以确保注释保真度。数据集分为训练集、测试集A和测试集B,总计110,507个场景和198,702个查询。它涵盖157个类别,总注释对象约730万个,每个图像对象数在1到250之间。数据集由上海交通大学人工智能学院的Chang Liu、Haoning Wu和Weidi Xie策划,采用Apache-2.0许可证。

KubriCount is a large-scale synthetic benchmark dataset for multi-granularity visual counting, proposed in the paper *Count Anything at Any Granularity* (Liu, Wu & Xie, SJTU 2026). It redefines open-world counting as a prompt-following task across five distinct semantic granularity levels, and stands as the most comprehensively annotated counting dataset released to date. This dataset contains 6,736 samples and is hosted via the FiftyOne platform. Each scene is a 1024×1024 synthetic image, generated via a four-stage automated pipeline: controllable 3D rendering using Kubric and Blender, mask-based image editing to narrow the sim-to-real gap, and quality filtering based on vision-language models to ensure annotation fidelity. The dataset is split into training set, test set A, and test set B, with a total of 110,507 scenes and 198,702 queries. It covers 157 categories, with approximately 7.3 million annotated objects in total, and the number of objects per image ranges from 1 to 250. The dataset was curated by Chang Liu, Haoning Wu, and Weidi Xie from the School of Artificial Intelligence, Shanghai Jiao Tong University, and is released under the Apache-2.0 license.

提供机构:
Voxel51
二维码
社区交流群
二维码
科研交流群
商业服务