From Registration to Manipulation: Evaluating Point Cloud Registration for Robotic Pick-and-Place
收藏资源简介:
From Registration to Manipulation: Offline and Robotic Experiment Results Description This archive contains the result data associated with: Luca Consoli, Andrea Ferrari, and Marco Carricato, "From Registration to Manipulation: Evaluating Point Cloud Registration for Robotic Pick-and-Place," ASME Letters in Translational Robotics, 2026. It includes the results of the offline evaluation of five point-cloud registration pipelines on GraspClutter6D and the records of robotic pick-and-place experiments performed with PPF and GMCNet. Object models, predefined grasp poses, and ChArUco ground-truth data used in the robotic experiments are also provided. Folder structure . |-- README.md |-- Offline_Evaluation/ | |-- graspclutter6d_extra/ | | `-- noshelves_scenes.json | `-- results/ | `-- <method>/ | |-- results.csv | |-- timing_stats.json | `-- parameters.txt `-- Robotic_Experiments/ |-- charuco/ | |-- charuco_poses.yaml | |-- ChArUco_5x5_40mm.pdf | `-- obj_<id>_charuco.stl |-- models/ | |-- obj_<id>.obj | `-- grasp_poses/ | `-- obj_<id>_grasp_transforms.json `-- results/ `-- <method>/<object_id>/<4D_or_6D>/<timestamped_run>/ Offline evaluation The evaluated methods are: PPF; FPFH + TEASER++; 3DSmoothNet + TEASER++; GMCNet; OMNet. Each pipeline was evaluated on all 1,000 GraspClutter6D scenes, 13 viewpoints, four RGB-D sensors, and 200 object models, corresponding to 758,940 object-instance attempts per method. The sensors are the Intel RealSense D415, Intel RealSense D435, Azure Kinect, and Zivid One+ M. Only visible/modal instance masks are used. An input is classified as invalid_input when fewer than 50 valid 3D points remain after object extraction. A valid input is classified as method_failed when the global method does not return a usable pose; otherwise its status is success. The ICP outcome is recorded separately. Every successful global pose is refined with the same point-to-point ICP: correspondence thresholds: 0.05, 0.02, 0.01, and 0.005 m; maximum iterations: 1,000 per threshold. Files provided for each method results.csv The main result table contains one row per registration attempt. It includes: scene, sensor, viewpoint, object, and instance identifiers; translation, rotation, ADD, and ADD-S errors before and after ICP; input, method, ICP, and total execution times; visibility and method-dependent point or correspondence counts; global and ICP statuses and failure reasons; model-to-camera poses before and after ICP. Pose fields follow: p_camera = R_m2c * p_model + t_m2c Translations and distance errors are in metres, rotation errors are in degrees, and times are in seconds. Visibility is the ratio between the non-zero pixels in the modal and amodal masks. Rows without a usable pose can contain empty metric and pose fields. For timing: total_time_s = method_time_s + icp_time_s input_time_s is stored separately and is not included in total_time_s. One-time preparation, training, initialization, and GPU warm-up are excluded from per-attempt timing. timing_stats.json Contains summary statistics for the timing values. The per-attempt columns in results.csv should be used to recompute statistics or analyze individual sensors and scene groups. parameters.txt Contains the main settings used by the corresponding method, including preprocessing, global registration, common ICP, and timing boundaries. Scene groups graspclutter6d_extra/noshelves_scenes.json lists the 691 non-shelf scenes, comprising bin and tabletop configurations. The remaining 309 scenes are shelf scenes. This auxiliary list is included because this grouping is not provided directly by GraspClutter6D. When analyzing the tables, use all rows for reliability statistics and only successful rows with finite values for pose accuracy. Join results across methods using the dataset identifiers rather than row order. Robotic experiments The experiments use a UR10e robot, a suction end-effector, and an eye-in-hand Intel RealSense D435. Eight objects are tested: 25, 35, 55, 101, 165, 174, 176, 180 For each object and method, three trials use a restricted 4-DOF displacement and three use a general 6-DOF displacement. The archive contains 48 trials for PPF, 48 for GMCNet, and 96 trials in total. Method Completed No feasible grasp Grasp-execution failure Total PPF 25 8 15 48 GMCNet 10 22 16 48 Models and grasp poses Robotic_Experiments/models/ contains the eight object-only OBJ models used for registration and pose-error evaluation. models/grasp_poses/ contains one JSON file per object, named: obj_<id>_grasp_transforms.json Each file stores the predefined suction-grasp frames with respect to the object's CAD reference frame (REF_CSYS). For every SUC_<n> grasp, the file provides its CAD step identifier, translation, rotation matrix, quaternion in xyzw order, and homogeneous transform T_ref_i. The translations are in millimetres. T_ref_i maps the suction-frame coordinates into the object reference frame. ChArUco files Robotic_Experiments/charuco/ contains: the printable ChArUco board; the object-and-marker-support STL assemblies; charuco_poses.yaml, containing the two fixed marker/object transforms for each tested object. In charuco_poses.yaml, each object ID has entries 0 and 1 for its two ChArUco boards. Each entry contains the translation in millimetres and the four-component rotation value used by the experimental pipeline to construct the corresponding fixed marker/object transform. The ChArUco supports are used only for independent ground-truth acquisition. They are excluded from the point cloud supplied to registration and from the model points used for ADD and ADD-S. Trial results Robotic_Experiments/results/<method>/<object_id>/<4D_or_6D>/ pose_place_run_<timestamp>/ results.csv trial_0000_obj_<object_id>/ Each timestamped directory corresponds to one physical trial. The detailed trial directory contains: trial_config.json: trial and acquisition settings; trial_events.jsonl: chronological event log; trial_results.json: outcome, errors, timing, selected grasp, and failure information; T_plate_object_*.json: estimated, initial, target, and available final object poses; SUC_*_T_*.json: selected grasp and pick/place transforms; estimation_scene/: RGB-D data, masks, and the point cloud used for registration; gt_initial/ and, when available, gt_final/: ChArUco ground-truth data. Files related to later manipulation stages can be absent when a trial terminates earlier. Use trial_results.json and trial_events.jsonl to determine the last completed stage. Transform convention A name of the form T_A_B maps coordinates from frame B to frame A: p_A = T_A_B * p_B T_plate_object = T_plate_cam * T_cam_charuco * T_charuco_object Initial and final ground-truth poses are computed from three images per trial. The target pose is computed once per object from ten images. Translations are averaged directly, while quaternion signs are aligned before averaging. Runtime translations and depth values are in metres unless otherwise specified. Grasp-pose and fixed CAD-derived ChArUco translations are stored in millimetres and converted to metres by the experimental pipeline. External data The original GraspClutter6D data are available from: Project: https://sites.google.com/view/graspclutter6d Dataset: https://huggingface.co/datasets/GraspClutter6D/GraspClutter6D API: https://github.com/SeungBack/graspclutter6dAPI Trained network weights are not redistributed. Their filenames are recorded in the corresponding parameters.txt files. Citation and contact Please cite this Zenodo record, the related paper, and the original GraspClutter6D publication when using these results. Zenodo DOI: 10.5281/zenodo.22874374 Paper DOI: not published yet The archive license is specified in the Zenodo record. External data and software remain subject to their original licenses. Contact: Luca Consoli, luca.consoli3@unibo.it.



