Neuromorphic Object Dataset for Detection and Reach: DVS128 Recordings of Seven Everyday Objects with Bounding-Box Annotations
收藏资源简介:
Neuromorphic Object Dataset for Detection and Reach — DVS128, 7 Everyday Objects Event-based (neuromorphic) recordings of seven everyday objects captured with a DVS128 dynamic vision sensor mounted on the end effector of a WidowX robotic arm, with manually annotated bounding boxes. Recorded at the Biomedical Engineering Lab, Federal University of Uberlândia, Brazil. Contents Item Description Raw event streams One recording per object, as produced by the DVS128 Dense representations Event surfaces obtained by integrating the raw streams over 100 ms windows Labels One bounding box per surface, annotated manually Acquisition Sensor: DVS128 dynamic vision sensor (128 × 128 pixels), attached to the side of the gripper of a WidowX Robot Arm Mark II. Motion: the arm followed one fixed programmed trajectory per recording, keeping the object inside the field of view while the relative motion between camera and object varied along the path. This yields several distinct event patterns for the same object. Duration: 24 s per recording. Scene: recorded against a complex background (laboratory furniture and equipment), not an isolated or blanked-out background. Classes Seven objects chosen for their presence in everyday activities: banana, cup, fork, key, knife, mug, orange. Dense representation Each raw recording was converted into surfaces of events: the value of a pixel is the sum of the polarities of the events that fell on it within the integration window, with P ∈ {−1, +1}. Positive polarity raises the value above the neutral mid-level, negative polarity lowers it. Integration window: 100 ms Surfaces per class: 208 Total surfaces: 1456 (The reach experiments in the associated paper used a 30 ms window at run time; 100 ms is the window used to build this dataset.) Label format Bounding boxes were annotated manually for every class, one box per surface, with five fields: centroid x, centroid y, box width, box height, class. Known limitations Spatial resolution is 128 × 128, the native resolution of the DVS128. Each class was recorded along a single programmed trajectory, so the diversity of event patterns per class is bounded by that trajectory. The key is the smallest object in the set and its event cluster is easily confused with sensor noise; in the detector reported in the associated paper it was the worst class (AP@0.5 = 0.45), and the knife was frequently predicted as fork. All recordings contain background events; there are no blank-background recordings. Related work Associated manuscript: Real-Time Neuromorphic Object Detection and Tracking with Saccade-Driven Event-Based Vision and YOLO (in preparation). Earlier work by the same group using this material or its method: E. B. Gouveia, E. L. S. Gouveia, V. T. Costa, A. Nakagawa-Silva, A. B. Soares, "Classification of Objects Using Neuromorphic Camera and Convolutional Neural Networks," Congresso Brasileiro de Engenharia Biomédica, 2020. E. B. Gouveia, L. V. Costa, E. L. S. Gouveia, V. T. Costa, A. Nakagawa-Silva, A. B. Soares, "An Object Tracking Using a Neuromorphic System Based on Standard RGB Cameras," Congresso Brasileiro de Engenharia Biomédica, 2020. E. L. S. Gouveia, E. B. Gouveia, A. Nakagawa-Silva, A. B. Soares, "Neuromorphic Vision-aided Semi-autonomous System for Prosthesis Control," Congresso Brasileiro de Engenharia Biomédica, 2020. E. B. Gouveia, G. F. Tavares, L. L. Almada, A. Nakagawa-Silva, E. L. S. Gouveia, M. J. Cunha, E. A. L. Junior, A. B. Soares, "Object detection using sparse data representation with convolutional neural networks for event-based cameras," Simpósio de Engenharia Biomédica, 2021.



