OpenAoE-2000h
收藏资源简介:
# Open-AoE — Egocentric Hand Manipulation Dataset ## Release Roadmap | Tier | Duration | Status | Tag | Samples | |------|----------|--------|-----|---------| | **nano** | 3 h | ✅ Released (batch 1) | `nano` | 111 | | **tiny** | 100 h | ✅ Released (batch 2) | `tiny` | 2710 | | **full** | 2000 h | 📋 Long-term goal (by end of Jul 2026) | — | TBD | All tiers are published as progressive releases in the **same** repos. Sample directories remain flat at the repo root. Nano samples use `aoe_*` IDs; tiny (batch 2) samples use `raw_<id>_seg_<n>` IDs (traceable to delivery segment ids). --- ## 📋 Overview This dataset contains **3D hand reconstruction data**, **camera trajectories**, and **fine-grained atomic action annotations** captured from an egocentric (first-person) viewpoint. The data originates from real-world smartphone-captured videos. After processing with an in-house pipeline (visual-SLAM-based camera trajectory estimation + parametric hand reconstruction), it provides complete MANO 3D hand pose parameters, metric-scale camera trajectories, and fine-grained atomic action annotations. ### Data Specifications | Item | Value | | --- | --- | | **Frame Rate** | 30 FPS | | **Video Resolution** | 1920×1080 / 1080×720 | | **Hand Model** | MANO (15 finger joint poses + 10-dim shape parameters) | | **Camera Trajectory** | Visual-SLAM-based camera trajectory estimation (metric-scale, camera→world homogeneous transform output) | | **Hand Reconstruction** | MANO parametric hand reconstruction (dual output: world + camera coordinate systems) | | **Annotation** | Automatic annotation (segment-level atomic actions, including scene and left/right hand attribution) | --- ## 📁 Directory Structure The standard output directory layout of one data item (one video clip) is shown below: ``` <sample>/ # Clip root directory ├── raw_video.mp4 # Raw captured video (1920×1080, 30 FPS) ├── video_info.json # Raw camera parameters and device info ├── ego_annotation/ # Annotation artifacts │ └── ego_action_annotation.json # Atomic action annotations (segment-level) └── ego_process/ # Processing artifacts ├── ego_hands_reconstruction/ # Hand 3D reconstruction results │ ├── hands.npz # Hand MANO parameters + camera poses (core data) │ ├── camera_traj.npz # Camera trajectory (visual SLAM) │ └── visualization/ # Visualization results │ ├── hands_combined.mp4 # Hand mesh overlay video │ └── overview.png # Trajectory overview image └── ego_undistorted_video/ # Undistorted video ├── raw_video_undistorted.mp4 # Undistorted video file └── undistorted_video_info.json # Undistorted camera parameters ``` --- ## 📄 Core Data File Details ### 1. `video_info.json` — Raw Video Camera Parameters Located in each sample's root directory. Records the original capture device information and camera intrinsics (calibrated values, with distortion). **Example**: ```json { "deviceInfo": { "brand": "Open-AoE", "model": "Open-AoE Pro", "androidVersion": "13" }, "cameraParams": { "usedCameraId": "2", "resolution": "1920x1080", "fx_pixels": 788.0, "fy_pixels": 591.0, "cx_pixels": 961.0, "cy_pixels": 539.0, "max_fx_pixels": 788.0, "max_fy_pixels": 788.0, "minFocusDistance_meters": 10, "lensDistortion": "[-0.002, 0.015, -0.009, 0.0, 0.0]" } } ``` | Field | Type | Description | | --- | --- | --- | | `deviceInfo.brand` | string | Smartphone brand | | `deviceInfo.model` | string | Smartphone model | | `deviceInfo.androidVersion` | string | Android version | | `cameraParams.usedCameraId` | string | Camera ID used for capture | | `cameraParams.resolution` | string | Video resolution `"width x height"` | | `cameraParams.fx_pixels` | float | Focal length fx (pixels, calibrated value) | | `cameraParams.fy_pixels` | float | Focal length fy (pixels, calibrated value) | | `cameraParams.cx_pixels` | float | Principal point cx (pixels) | | `cameraParams.cy_pixels` | float | Principal point cy (pixels) | | `cameraParams.max_fx_pixels` | float | Maximum focal length fx (pixels) | | `cameraParams.max_fy_pixels` | float | Maximum focal length fy (pixels) | | `cameraParams.minFocusDistance_meters` | float | Minimum focus distance (meters) | | `cameraParams.lensDistortion` | string | Lens distortion coefficients (JSON-formatted array, 5 coefficients `[k1, k2, k3, p1, p2]`) | > **Note**: The coefficient order `[k1, k2, k3, p1, p2]` corresponds to the `lensDistortionOrder` field (present in most samples, consistently `"k1,k2,k3,p1,p2"` when available). If `lensDistortionOrder` is absent for a given sample, `lensDistortion` defaults to zero distortion, i.e. `[0.0, 0.0, 0.0, 0.0, 0.0]`. > **Note**: In the raw intrinsics, `fx_pixels` and `fy_pixels` are not equal (different horizontal/vertical sampling ratios). When working with hand data and camera trajectories, use the undistorted intrinsics (see `undistorted_video_info.json` below). --- ### 2. `undistorted_video_info.json` — Undistorted Camera Parameters Located in the `ego_process/ego_undistorted_video/` directory. **Use the parameters from this file when working with hand data and camera trajectories.** Based on `video_info.json`, the distortion coefficients are zeroed out, the intrinsics are recomputed, and additional fields are added: field of view (FOV), video filename, and frame rate. **Example** (from an example sample): ```json { "deviceInfo": { "brand": "Open-AoE", "model": "Open-AoE Pro", "androidVersion": "13" }, "cameraParams": { "usedCameraId": "2", "resolution": "1920x1080", "fx_pixels": 788.0, "fy_pixels": 791.0, "cx_pixels": 962.0, "cy_pixels": 540.0, "max_fx_pixels": 788.0, "max_fy_pixels": 788.0, "minFocusDistance_meters": 10, "lensDistortion": "[0.0, 0.0, 0.0, 0.0, 0.0]", "fov_x_degrees": 101.3, "fov_y_degrees": 68.6, "fov_x_radians": 1.77, "fov_y_radians": 1.20 }, "video_filename": "raw_video_undistorted.mp4", "fps": 30 } ``` Key differences from `video_info.json`: | Difference | Raw Intrinsics File | Undistorted Intrinsics File | Description | | --- | --- | --- | --- | | `cameraParams.fx_pixels` / `fy_pixels` | Raw calibrated values (fx≠fy) | Recomputed after undistortion | Used for projecting hand / trajectory data | | `cameraParams.cx_pixels` / `cy_pixels` | Raw values | Recomputed after undistortion | — | | `cameraParams.lensDistortion` | Raw distortion coefficients | `"[0.0, 0.0, 0.0, 0.0, 0.0]"` | Zero after undistortion | | `cameraParams.fov_x_degrees` / `fov_y_degrees` | ❌ Absent | ✅ Present | Horizontal / vertical FOV (degrees) | | `cameraParams.fov_x_radians` / `fov_y_radians` | ❌ Absent | ✅ Present | Horizontal / vertical FOV (radians) | | `video_filename` | ❌ Absent | ✅ Present | Undistorted video filename (`raw_video_undistorted.mp4`) | | `fps` | ❌ Absent (also absent in raw file) | ✅ Present | Video frame rate (30) | --- ### 3. `hands.npz` — Hand 3D Reconstruction Data (Core Data) Located in the `ego_process/ego_hands_reconstruction/` directory, generated by the parametric hand reconstruction module. Represented using the [MANO](https://mano.is.tue.mpg.de/) parametric hand model. > **Notation**: In the shapes below, `T` = number of video frames, `2` = the two-hands dimension (index `0` = left hand, `1` = right hand). **Loading:** ```python import numpy as np data = np.load("hands.npz") print(data.files) # ['R_w2c', 't_w2c', 'R_c2w', 't_c2w', 'pred_trans', 'pred_rot', # 'pred_trans_cam', 'pred_rot_cam', 'pred_hand_pose', 'pred_betas', # 'pred_valid', 'focal'] ``` #### 3.1 Camera Pose Parameters | Field | Shape | Dtype | Unit | Description | | --- | --- | --- | --- | --- | | `R_w2c` | `(T, 3, 3)` | float32 | — | World→camera rotation matrix | | `t_w2c` | `(T, 3)` | float32 | meters | World→camera translation vector | | `R_c2w` | `(T, 3, 3)` | float32 | — | Camera→world rotation matrix | | `t_c2w` | `(T, 3)` | float32 | meters | Camera→world translation vector | | `focal` | `()` | float64 | pixels | Focal length scalar (consistent with `fx_pixels` in `undistorted_video_info.json` and `intrinsic[0,0]`) | **Coordinate transform relations:** ```python # World coordinate system → camera coordinate system p_cam = R_w2c[t] @ p_world + t_w2c[t] # Camera coordinate system → world coordinate system p_world = R_c2w[t] @ p_cam + t_c2w[t] ``` #### 3.2 Hand MANO Parameters (World Coordinate System) Derived from the camera-coordinate results via the camera pose: `p_world = R_c2w @ p_cam + t_c2w`. | Field | Shape | Dtype | Unit | Description | | --- | --- | --- | --- | --- | | `pred_trans` | `(2, T, 3)` | float32 | meters | Wrist translation (world coordinate system) | | `pred_rot` | `(2, T, 3)` | float32 | radians | Wrist global rotation (axis-angle, world coordinate system) | #### 3.3 Hand MANO Parameters (Camera Coordinate System) The native output of hand reconstruction in the camera coordinate system (each frame relative to that frame's camera). Conversion to/from the world coordinate system: `p_cam = R_w2c @ p_world + t_w2c`. | Field | Shape | Dtype | Unit | Description | | --- | --- | --- | --- | --- | | `pred_trans_cam` | `(2, T, 3)` | float32 | meters | Wrist translation (camera coordinate system) | | `pred_rot_cam` | `(2, T, 3)` | float32 | radians | Wrist global rotation (axis-angle, camera coordinate system) | #### 3.4 MANO Shape and Pose Parameters | Field | Shape | Dtype | Unit | Description | | --- | --- | --- | --- | --- | | `pred_hand_pose` | `(2, T, 45)` | float32 | radians | 15 finger joint poses (3-dim axis-angle each), coordinate-system independent | | `pred_betas` | `(2, T, 10)` | float32 | — | MANO shape parameters (10-dim PCA coefficients), coordinate-system independent | #### 3.5 Validity Mask | Field | Shape | Dtype | Unit | Description | | --- | --- | --- | --- | --- | | `pred_valid` | `(2, T)` | float32 | — | Per-frame detection validity, valued `1.0` (valid) / `0.0` (invalid). Convert to bool before use | > **⚠️ Important**: For all arrays with shape `(2, ...)`, the first-dimension index `[0]` = **left hand**, `[1]` = **right hand**. `pred_valid` is a float mask (`0.0`/`1.0`) and must be converted to bool before use; for invalid frames (`pred_valid == 0`), the values of other fields are unreliable and must be filtered out first. #### Usage Example ```python import numpy as np from scipy.spatial.transform import Rotation data = np.load("hands.npz") # Get world-coordinate positions for valid frames of the right hand (index 1) valid_right = data['pred_valid'][1] > 0 # (T,) validity mask to bool, [1]=right hand right_trans = data['pred_trans'][1] # (T, 3) wrist translation, world coords, meters right_trans_valid = right_trans[valid_right] # (N_valid, 3) filtered valid frames # Project a world-coordinate hand position to the image focal = data['focal'] # () focal length scalar, pixels (= fx) R_w2c = data['R_w2c'] # (T, 3, 3) world→camera rotation matrix t_w2c = data['t_w2c'] # (T, 3) world→camera translation vector, meters frame_idx = 100 p_cam = R_w2c[frame_idx] @ right_trans[frame_idx] + t_w2c[frame_idx] # fx/fy/cx/cy come from undistorted_video_info.json or the intrinsic in camera_traj.npz u = fx * p_cam[0] / p_cam[2] + cx v = fy * p_cam[1] / p_cam[2] + cy # Reconstruct the hand mesh using the MANO model (right hand = index 1) rot_aa = data['pred_rot'][1, frame_idx] # (3,) wrist rotation, axis-angle, radians hand_pose = data['pred_hand_pose'][1, frame_idx] # (45,) joint poses, radians betas = data['pred_betas'][1, frame_idx] # (10,) shape parameters, dimensionless ``` --- ### 4. `camera_traj.npz` — Camera Trajectory Located in the `ego_process/ego_hands_reconstruction/` directory. Generated by the visual-SLAM-based camera trajectory estimation module. **Loading:** ```python import numpy as np data = np.load("camera_traj.npz") print(data.files) # ['cam_c2w', 'intrinsic'] ``` | Field | Shape | Dtype | Unit | Description | | --- | --- | --- | --- | --- | | `cam_c2w` | `(T, 4, 4)` | float64 | meters (translation part) | Per-frame camera-to-world 4×4 homogeneous transform matrix | | `intrinsic` | `(3, 3)` | float64 | pixels | Camera intrinsic matrix (undistorted, consistent with `undistorted_video_info.json`) | > **Note**: In this dataset version, `camera_traj.npz` **contains only** the `cam_c2w` and `intrinsic` fields; it **does not** include per-frame depth maps (`depths`). #### Transform Matrix `cam_c2w[t]` is the 4×4 homogeneous transform matrix for frame t, which transforms points from the camera coordinate system to the world coordinate system: ``` cam_c2w = [[R(3×3), t(3×1)], [0 0 0, 1 ]] p_world = cam_c2w @ [p_cam; 1] ``` #### Intrinsic Matrix ``` intrinsic = [[fx, 0, cx], [ 0, fy, cy], [ 0, 0, 1]] ``` **Concrete example** (from an example sample, 1920×1080 resolution): ``` [[788.0 0.0 962.0] [ 0.0 791.0 540.0] [ 0.0 0.0 1.0]] ``` #### Projecting a 3D World Point to the Image ```python import numpy as np data = np.load("camera_traj.npz") cam_c2w = data['cam_c2w'] # (T, 4, 4) camera-to-world homogeneous transform matrix K = data['intrinsic'] # (3, 3) intrinsic matrix, pixels # World point → camera frame → image T_w2c = np.linalg.inv(cam_c2w[frame_idx]) # world-to-camera p_cam = T_w2c @ np.array([x, y, z, 1.0]) # world coords → camera coords p_img = K @ p_cam[:3] # camera coords → image coords u, v = p_img[0] / p_img[2], p_img[1] / p_img[2] # normalize to pixel coords ``` --- ### 5. `ego_action_annotation.json` — Atomic Action Annotations Located in the `ego_annotation/` directory. Contains **segment-level** fine-grained atomic manipulation annotations. It is a JSON array in which each element is a time segment. The number of segments varies per sample (e.g., 7 ~ 39 segments). **Annotation flow**: video → atomic action recognition → temporal segmentation → JSON output #### 5.1 Annotation Format ```json [ { "id": 1, "start_ts": "0.0", "end_ts": "3.0", "start_frame": "0", "end_frame": "90", "scene": "bedroom", "atomic_action": [ { "verb": "grasp", "object": "plastic lid", "hand": "right", "description": "The person grasps a small clear plastic lid from the table surface.", "bbox": [630, 320, 720, 420], "confidence": 0.95 } ] } ] ``` #### 5.2 Field Descriptions | Field | Type | Description | | --- | --- | --- | | `id` | int | Annotation segment index (incrementing from 1) | | `start_ts` | string | Start timestamp (seconds) | | `end_ts` | string | End timestamp (seconds) | | `start_frame` | string | Start frame number (≈ round(start_ts × fps), fps=30) | | `end_frame` | string | End frame number (≈ round(end_ts × fps), fps=30) | | `scene` | string | Current scene description (English, e.g., `"bedroom"`, `"laundry room"`) | | `atomic_action` | array | List of atomic actions (usually 1 action per segment) | | `atomic_action[].verb` | string | Action verb (English, fine-grained manipulation primitive) | | `atomic_action[].object` | string | Interacted object (English) | | `atomic_action[].hand` | string | Hand performing the action: `"left"` / `"right"` / `"both"` / `"none"` | | `atomic_action[].description` | string | Detailed action description (English, one sentence) | | `atomic_action[].bbox` | array | Action region bounding box `[x1, y1, x2, y2]`, in `raw_video.mp4` pixel coordinates (1920×1080) | | `atomic_action[].confidence` | float | Confidence score (0.0 ~ 1.0) | #### 5.3 Annotation Conventions - **Full coverage**: Annotations span from 0.0s to the last frame of the video, with no gaps or overlaps - **Temporal continuity**: `clip[i].end_ts == clip[i+1].start_ts` - **One action per segment**: Each time segment usually contains one atomic manipulation primitive - **Verb choice**: Use fine-grained manipulation verbs (e.g., `grasp`, `twist`, `slide`) and avoid vague verbs (e.g., `use`, `do`) - **Visual observability**: Only describe visually observable actions; do not infer hidden objects or actions > **Note**: > > - `start_ts`, `end_ts`, `start_frame`, and `end_frame` are all **string types** and require conversion to numeric values before use. > - Frame numbers / bboxes are in the original pixel / frame space of `raw_video.mp4`; to align with the undistorted video or the hand / trajectory data, index by frame number (the frame ordering is consistent). --- ### 6. Video Files | File | Location | Description | | --- | --- | --- | | `raw_video.mp4` | Sample root | Original captured video (1920×1080, 30 FPS) | | `raw_video_undistorted.mp4` | `ego_process/ego_undistorted_video/` | Undistorted video | | `hands_combined.mp4` | `ego_process/ego_hands_reconstruction/visualization/` | Visualization video with 3D hand meshes overlaid on video frames (left hand purple, right hand blue) | | `overview.png` | `ego_process/ego_hands_reconstruction/visualization/` | Overview image of camera trajectory and hand motion trajectory | --- ## 🔧 Coordinate System Definitions ### Camera Coordinate System (OpenCV Convention) ``` Z (forward / optical axis) ↑ | | ●──────→ X (right) / / Y (down) ``` - **X axis**: points right (image u direction) - **Y axis**: points down (image v direction) - **Z axis**: points forward (optical axis, depth positive) ### World Coordinate System Defined by visual SLAM, referenced to the first frame's camera pose. All 3D coordinates are in **meters**. ### Coordinate Transforms ```python # Camera coordinate system → world coordinate system p_world = R_c2w @ p_cam + t_c2w # 3×3 rotation + 3×1 translation in hands.npz p_world = cam_c2w @ [p_cam; 1] # 4×4 homogeneous transform in camera_traj.npz # World coordinate system → camera coordinate system p_cam = R_w2c @ p_world + t_w2c # 3×3 rotation + 3×1 translation in hands.npz # Camera coordinate system → image coordinate system (fx/fy/cx/cy from undistorted_video_info.json or the intrinsic matrix) # The focal scalar in hands.npz = fx (after undistortion fx and fy still differ slightly; prefer the intrinsic fx/fy for projection) u = fx * p_cam[0] / p_cam[2] + cx v = fy * p_cam[1] / p_cam[2] + cy ``` --- ## Key Notes ### Validity Checks - **Unit consistency**: All 3D coordinates are in **meters** - **Rotation representation**: Hand rotations use the **axis-angle** representation and must be converted to rotation matrices before use - **Hand index**: For arrays in `hands.npz` with shape `(2, ...)`, `[0]` = **left hand**, `[1]` = **right hand** - **Validity mask type**: `pred_valid` is **float32** (`0.0`/`1.0`) and must be converted to bool before use - **Fixed intrinsics**: Camera intrinsics remain constant across all frames (after undistortion) - **Validity filtering**: You must use `pred_valid` to filter out invalid frames; data for invalid frames is unreliable - **Timestamp type**: Timestamps and frame numbers in annotation files are all **strings** and require conversion before use --- ## 🔍 Visualization End-to-end dataset visualization is provided by the standalone tool **`aoe-visualization`** (located at `release/aoe-visualization/`). It re-renders MANO hand reconstruction results and overlays atomic action annotations onto the video, producing a full review video `AoE_output_vis.mp4` per sample (including: undistorted ego video + hand mesh/keypoint overlay, action-annotation info panel, world-coordinate 3D panel, and bottom timeline). ### Installation ```bash cd /path/to/release/aoe-visualization pip install -r requirements.txt # numpy, opencv-python, pyrender, trimesh, PyOpenGL ``` > High-quality hand mesh rendering requires an available EGL/OpenGL offscreen context (typically on a GPU machine). If no GL context can be created, it automatically falls back to pure NumPy rendering (lower quality). ### Usage ```bash cd /path/to/release/aoe-visualization # Visualize a single sample -> output/<sample_name>/AoE_output_vis.mp4 python visualize.py --sample /path/to/<sample_id> # Visualize an entire dataset directory python visualize.py --data_dir /path/to/dataset_dir --output_dir ./output ``` where `<sample_id>` is the root directory path of the video clip. See `aoe-visualization/README.md` for more details (rendering pipeline, MANO assets, dependencies, etc.). ### Visualization Outputs - `AoE_output_vis.mp4`: end-to-end review video generated by `aoe-visualization` (hand mesh/keypoint overlay + action annotations + world-coordinate 3D panel + timeline) Pre-rendered visualization files shipped with the dataset are located under each sample's `ego_process/ego_hands_reconstruction/visualization/` directory: - `hands_combined.mp4`: 3D hand mesh overlaid on video frames (left hand purple, right hand blue) - `overview.png`: overview image of the camera trajectory and hand motion trajectory



