RGBD-VideoCount
收藏资源简介:
RGBD-VideoCount数据集由哈尔滨工业大学(威海)等机构构建,是首个面向拥挤遮挡场景的RGB-D视频目标计数数据集。该数据集提供195个同步RGB-D视频片段,覆盖6个目标类别,并包含深度信息与多类别共存场景,数据规模达195个序列。数据集通过真实仓库货架等密集场景采集,并标注了目标边界框和深度信息。该数据集旨在推动基于深度信息的视频目标计数研究,解决RGB信息在严重遮挡和外观相似条件下的局限性,为自动化库存管理、货架巡检等实际应用提供标准化评估基准。
RGBD-VideoCount Dataset is constructed by Harbin Institute of Technology (Weihai) and other relevant institutions. It is the first RGB-D video object counting dataset targeting crowded and occluded scenes. The dataset provides 195 synchronized RGB-D video clips covering 6 object categories, includes depth information and multi-class co-occurrence scenarios, and has a total of 195 sequences. Collected from dense real-world scenarios such as warehouse shelves, the dataset is annotated with target bounding boxes and depth information. This dataset aims to advance research on depth-based video object counting, address the limitations of RGB information under severe occlusion and visually similar conditions, and provide a standardized evaluation benchmark for practical applications including automated inventory management and shelf inspection.
RGBD-VideoCount 数据集概述
基本信息
- 许可证: CC-BY-4.0(知识共享署名4.0国际许可协议)
- 任务类型: 目标检测(object detection)
- 标签: RGB-D、视频目标计数、拥挤场景、遮挡
数据集简介
RGBD-VideoCount 是一个用于拥挤和遮挡场景下视频目标计数的 RGB-D 视频数据集。该数据集提供同步的 RGB 帧和深度图,以及用于评估检测、跨帧关联和视频级去重的实例级标注。
数据集规模
- 视频片段数量: 195 个 RGB-D 视频片段
- 目标类别: 6 个类别
- 精细标注帧数: 2,032 帧
- 实例边界框数量: 77,638 个
- 场景特点: 多类别货架和拥挤目标场景
数据内容
数据集包含以下组成部分:
- RGB 视频帧
- 与 RGB 帧对齐的深度图
- 实例级边界框标注
- 视频级计数标注
- 数据划分(训练/验证/测试)
- 可视化示例(exemplars)
目录结构
RGBD-VideoCount/ |- images/ # RGB视频帧 |- Depth_Data/ # 深度图(与RGB帧对齐) |- object_annotations/ # 实例级边界框标注 |- count_annotations/ # 视频级计数标注 |- dataset_split.json # 训练、验证、测试划分 |- video_class.txt # 类别元数据 |- exemplars_train.json # 训练集可视化示例 |- exemplars_val.json # 验证集可视化示例 `- exemplars_test.json # 测试集可视化示例
引用信息
该数据集对应的论文为:
- 论文标题: Depth-Guided Video Object Counting in Crowded Scenes
- 发表于: Proceedings of the 34th ACM International Conference on Multimedia(2026年)
- DOI: 10.1145/3767308.3835482
局限性说明
- 数据集主要聚焦于拥挤目标场景,可能无法代表所有真实世界环境
- 性能可能受到深度质量、严重外观模糊、相机运动以及未见目标类别的影响
- 用户在实际部署前需自行评估适用性
相关资源
- 模型权重: https://huggingface.co/aerospace123/DG-Net
- 数据集页面: https://huggingface.co/datasets/aerospace123/RGBD-VideoCount
- 代码仓库: https://github.com/streamer-AP/DG-Net

- 1Depth-Guided Video Object Counting in Crowded Scenes哈尔滨工业大学(威海); 香港城市大学; 中国科学技术大学; 哈尔滨工业大学青岛研究院 · 2026年




