遇见数据集

UrbanEgo: A Multimodal First-Person Urban Perception Dataset

收藏
Zenodo2026-09-30 更新2026-10-01 收录
官方服务:

资源简介:

What this dataset is This dataset captures the world from a walking pedestrian's point of view. A researcher wearing a Microsoft HoloLens 2 headset and a small backpack of recording hardware walked through the city of Aveiro, Portugal, for eight runs, with asynchronous sensor streams logged using a shared host clock. The collection covers one wearer, daytime, and fair weather between May and July 2026. Each recording ("run") contains: RGB video (what the wearer saw); depth frames (distance to nearby surfaces) and matching infrared intensity images; the head pose and heading (where the head was and which way it faced); an IMU stream (head tilt: yaw, pitch, roll); two independent GPS tracks (the wearer's position on the map); and an offline object-detection layer (people, bicycles, and vehicles found in the video). How it was gathered The data was originally recorded in a compact binary format on the device and then converted, offline, into the standard, widely readable formats you see here (MP4, PNG, JSON Lines, CSV). There are eight runs, one folder each, named runN_YYYYMMDD_HHMMSS where N is the chronological order (1–8) and the rest identifies the session start date and time in UTC. Recording runs The eight recording sessions span approximately 2 h 14 min across the released sensor streams; the exported RGB videos total 2 h 08 min 11 s. The run folders occupy approximately 7.8 GB. Session duration is the interval from the earliest sensor timestamp to the latest sensor timestamp in a run. RGB duration is the exported MP4 video-track duration. These differ because streams start and stop at different times, including GPS recording after video ends. Durations below are rounded to the nearest second; sizes use decimal GB (10⁹ bytes), including each run's metadata. Run Date Session start (UTC) Session duration RGB duration Size (GB) Route 1 2026-05-06 15:26:21 29 min 45 s 29 min 39 s 1.879 University campus loop 2 2026-05-06 16:39:35 11 min 08 s 11 min 04 s 0.660 Rua da Pêga, lakeside arterial 3 2026-05-07 15:19:20 17 min 02 s 16 min 57 s 1.098 Campus and hospital roundabout 4 2026-07-01 12:47:34 11 min 08 s 11 min 04 s 0.626 Repeat of Run 2 5 2026-07-13 17:35:11 20 min 01 s 18 min 13 s 1.144 Repeat of Run 3 6 2026-07-14 16:15:37 18 min 48 s 14 min 58 s 0.848 Ponte dos Botirões 7 2026-07-14 17:14:27 16 min 03 s 15 min 58 s 0.915 Rossio, beside the canal 8 2026-07-14 17:41:38 10 min 23 s 10 min 19 s 0.615 Praça do Peixe Runs 1, 3, and 5 cover the University of Aveiro campus and neighbouring streets; Runs 2 and 4 follow Rua da Pêga; Runs 6–8 cover the touristic city centre. Runs 7 and 8 contain the highest pedestrian densities; Runs 3, 4, and 6 contain the highest vehicle densities. Run 1 combines moderate pedestrian and vehicle activity with the most cycling activity. Synchronization and timestamp fields Sensor JSON Lines records carry ts_unix_ns, a host wall-clock timestamp in nanoseconds since the Unix epoch (1970-01-01 UTC), and ts_mono_ns, a host monotonic timestamp in nanoseconds. The streams are asynchronous and have different rates: exported MP4 playback is 20 fps, retained RGB metadata varies by run (approximately 19–24 samples/s), depth is approximately 5 Hz, GPS is approximately 0.8 Hz, and heading/orientation is approximately 13 Hz. Use nearest ts_unix_ns samples to associate sensor streams within their overlapping coverage, retaining the time difference and rejecting matches across large gaps. The YOLO files use ts_ns for the same wall-clock timestamps carried by the corresponding RGB metadata rows. They contain frame and t, and do not contain the common sensor fields seq, ts_unix_ns, or ts_mono_ns. Layout of one run runN_YYYYMMDD_HHMMSS/ manifest.json run summary: identifiers, per-stream counts, file layout rgb.mp4 egocentric video (H.264) + audio (AAC), constant 20 fps, real time rgb_frames.jsonl retained camera-frame sequence: timestamp, pose, camera intrinsics depth/000001.png … depth image, 16-bit PNG, value = millimetres to the surface ab/000001.png … active-brightness (infrared) image, 16-bit PNG, same grid as depth depth_frames.jsonl one line per depth frame: timestamp and head pose calibration/ depth camera calibration + the per-pixel ray table (see below) gps_receiver.jsonl Receiver GPS track from the external receiver connected to the Jetson via USB gps_phone.jsonl Phone GPS track from the OwnTracks phone app heading.jsonl head heading in degrees from North imu.jsonl head orientation: yaw, pitch, roll (degrees) yolo/ offline object detections (detections.jsonl, frames.jsonl, meta.json) The egocentric video: rgb.mp4 and rgb_frames.jsonl rgb.mp4 is H.264 video (1280×720, constant 20 fps) with AAC stereo audio (48 kHz). The camera stream was re-encoded and retimed for playback. Its frames do not correspond one-to-one to the retained camera metadata or YOLO frame rows. Use rgb_frames.jsonl for recorded camera timestamps and pose metadata. For browsing, first_rgb_metadata_ts_unix_ns + playback_seconds × 1e9 provides an approximate timestamp lookup; seek by playback time and choose the nearest metadata timestamp. Faces and vehicle licence plates in the released RGB video were processed with automatic detection and blurring. The process can miss identifiable content; residual identifiable content may remain. rgb_frames.jsonl lists the retained camera-frame sequence used by the detection pipeline, at a variable rate, one record per retained frame. Its compact frame index joins directly to yolo/frames.jsonl and yolo/detections.jsonl. The original acquisition count is preserved separately as manifest.json → stats.rgb_source_frames. Key Meaning frame Compact retained-camera metadata index (1-based); also the YOLO frame index. It is not an MP4 frame index. source_timestamp Recorded HoloLens device-clock value. Its clock conversion is not supplied; do not treat it as Unix time or interchange it with host timestamps. pose Recorded 4×4 local pose matrix, flattened in row-major order. The upper-left 3×3 block is rotation; the last row contains translation (x, y, z) in metres. See coordinate conventions below before applying it to camera-space points. pv_focal_length Camera focal length in pixels [fx, fy]. pv_principal_point Optical centre in pixels, [cx, cy]. pv_exposure_time Frame exposure time (device units). pv_iso_speed Camera ISO (sensor gain). pv_white_balance White-balance value (colour temperature). pv_resolution Frame size [width, height] = [1280, 720]. Depth and 3D: depth/, ab/, depth_frames.jsonl, calibration/ The HoloLens long-throw depth sensor produces two 320×288 images per frame: depth/NNNNNN.png — a 16-bit PNG where each pixel value is the distance in millimetres to the nearest surface along that pixel's ray. A value of 0 means "no return" (too far, too close, or no reflection). Depth is short-range (a few metres). ab/NNNNNN.png — the active brightness (infrared) image of the same instant on the same pixel grid; useful as a low-light greyscale view. depth_frames.jsonl has one line per depth frame with frame (index, 1-based), the usual timestamps, and the recorded local pose (same matrix storage as above; the source frame can differ from the RGB pose). The image depth/NNNNNN.png corresponds to the row whose frame is NNNNNN. Turning depth into 3D points. The depth is already lens-distortion-corrected, and the calibration/ folder ships a per-pixel unit-ray table so you can reconstruct 3D in one step, point = ray × distancein any language: calibration/ray_x.csv, ray_y.csv, ray_z.csv — three plain CSV grids, every 288 rows × 320 columns (one number per pixel): the x, y, z components of that pixel's unit ray. calibration/calibration.json — the metadata: Key Meaning sensor Sensor name (hololens2_rm_depth_longthrow). width, height Depth image size (320 × 288). depth_unit Unit of the PNG values (mm). depth_scale Divide PNG values by this to get metres (1000). rectified true: the depth is already distortion-corrected. intrinsics_4x4 Depth camera intrinsics (4×4). extrinsics_4x4 Recorded calibration transform (4×4). Its source/destination frame direction in this export requires confirmation before use. ray_files Names of the three ray CSVs. note One-line reconstruction reminder. Coordinate conventions The pose and calibration matrices refer to the local HoloLens tracking and sensor coordinate frames. Local translation is measured in metres and is not a WGS-84 position or a georeferenced object location. Row-major storage specifies how to reshape the 16 numbers. For a confirmed row-vector transform M, a homogeneous point is transformed as [x, y, z, 1] @ M, with translation in the last row. Position (two sources): gps_receiver.jsonl and gps_phone.jsonl Two independent GPS tracks, Receiver GPS and Phone GPS, are provided for redundancy (a satnav fix can drift or drop out). Receiver GPS (gps_receiver.jsonl) — from the external GPS receiver connected to the Jetson via USB. Key Meaning latitude, longitude Position in decimal degrees (WGS-84). altitude Not available for this source — always the sentinel 800001.0. Use the phone altitude instead. Some receiver rows carry unavailable-coordinate sentinels (latitude = 900000000, longitude = 1800000000, and altitude = 800001). Drop rows with abs(latitude) > 90 or abs(longitude) > 180. Apply coordinate-range checks to both GPS sources, then filter jumps that imply implausible walking speeds. Range checks alone do not detect every erroneous fix. The phone track in Run 3 contains particularly frequent outliers. Phone GPS (gps_phone.jsonl) — from the OwnTracks phone app. Key Meaning latitude, longitude Position in decimal degrees (WGS-84). altitude Altitude in metres (valid for this source). speed_mps Reported ground speed under the exported metres-per-second field. Optional; a missing value is unknown, not evidence of zero speed. Check implausible values before use. source_ts_ns Optional source-provided timing field found in Runs 4–8. Its values are stored in nanoseconds; the producer clock/epoch mapping remains unverified. Use ts_unix_ns for cross-stream association. Coverage note: the phone track is available from the start of each run; the receiver track can start a bit later (up to ~70 s in a couple of early runs while it acquired a fix). The GPS and video spans do not always coincide at the ends: usually the GPS stops a little before the video, but in some runs the GPS track keeps recording after the video ends. The phone track can bridge gaps in the receiver track where it has valid coverage. gps_receiver.jsonl can also contain the optional generation_delta_time field in Runs 4–8. This is preserved producer timing metadata; its units, reference epoch, and wrap behaviour have not been confirmed from the export code. It is not a replacement for ts_unix_ns. Where the head faced: heading.jsonl and imu.jsonl heading.jsonl — the compass direction the head faced. Key Meaning heading Degrees clockwise from North (0 = North, 90 = East), corrected per run using GPS course on moving straight segments and the facing direction visible in RGB. imu.jsonl — the head's orientation from the headset's inertial sensors. Key Meaning yaw Recorded orientation about the vertical axis (degrees). Use the corrected heading.jsonl stream for map bearing. pitch Up/down head tilt (degrees). roll Side-to-side head tilt (degrees). Offline object detections: yolo/ An automatic pass over the RGB video with a YOLO11x detector and a BoT-SORT tracker (re-identification), grouping road users into people, bicycles, and vehicles (car, motorcycle, bus, truck). Join YOLO frame to rgb_frames.jsonl.frame, or YOLO ts_ns to rgb_frames.jsonl.ts_unix_ns. These joins associate detections with the retained camera metadata sequence. yolo/detections.jsonl — one row per tracked object per frame. Key Meaning frame Frame index (matches rgb_frames.jsonl). t Seconds relative to the camera-processing origin, which can differ from session start and MP4 playback zero. Use ts_ns for sensor alignment. ts_ns Wall-clock time (nanoseconds since the Unix epoch). cls COCO class id (0 person, 1 bicycle, 2 car, 3 motorcycle, 5 bus, 7 truck). name Class name. id Tracker-assigned identifier within this run. Tracks can fragment or switch identity; IDs are not guaranteed to identify unique physical objects. conf Detection confidence, 0–1. box Pixel bounding box [x1, y1, x2, y2]. yolo/frames.jsonl contains one row per retained camera metadata frame, including frames with no detected objects. Including empty frames avoids excluding zero-detection observations from averages; the counts remain subject to detector errors. Key Meaning frame, t, ts_ns As above. n Total objects detected in the frame. counts Count per group: person, bicycle, vehicle. per_class Count per COCO class: person, bicycle, car, motorcycle, bus, truck. yolo/meta.json records the weight filename (yolo11x.pt), tracker configuration filename (botsort_reid.yaml), confidence threshold (0.25), IoU threshold (0.5), inference size (1280), processed-frame count, and camera-timestamp span. effective_fps describes the processed metadata sequence, not the exported MP4. manifest.json: the per-run summary Key Meaning run_name This folder's name (runN_YYYYMMDD_HHMMSS). source_folder, run_id Original recording folder and its date-time ID. run_number Chronological order, 1–8. format_version Metadata layout version (2 for this package), separate from the eventual Zenodo dataset release version. stats.rgb_source_frames Original acquisition count preserved from the source manifest. The original HLP2 files are not included. stats.rgb_metadata_frames Number of retained RGB metadata rows; YOLO frames use this sequence. stats.rgb_encoded_frames, stats.rgb_fps Actual exported MP4 video-track frame count and constant playback rate (20 fps). stats.rgb_metadata_fps Mean retained camera-metadata rate, computed as record count divided by its timestamp span. stats.rgb_duration_s Exported MP4 video-track duration, seconds. stats.duration_s is an alias for this value. stats.session_duration_s Span from the earliest sensor timestamp to the latest sensor timestamp, seconds. Other stats counts Number of depth frames, GPS, heading, and orientation samples in the release. timing Absolute session start/end, first RGB metadata offset, and each sensor stream's timestamp span. A span does not imply uninterrupted coverage. frame_alignment Association rules for original acquisition, retained metadata/YOLO, and exported MP4 frames. layout Which files make up each part of the run (rgb, depth, calibration, streams, detections).

提供机构:
Zenodo
创建时间:
2026-09-30
二维码
社区交流群
二维码
科研交流群
商业服务