TPD: Thermal Pose Dataset
收藏资源简介:
TPD (Thermal Pose Dataset) is a paired RGB and long-wave infrared (LWIR) human pose dataset for surveillance-oriented 2D human pose estimation. It contains 1,895 hardware-registered frame pairs across 12 activity classes, with 2,364 people annotated on a 17-point extended-MPII skeleton, for 40,188 keypoint coordinates in total. Existing thermal pose resources are largely confined to controlled indoor scenes, fixed subject-to-camera distances, and limited activity diversity. TPD was captured across institutional interiors (public libraries, cafes, gyms), open outdoor spaces under solar thermal loading (parks), and high-occlusion natural cover (dense foliage), at subject-to-camera distances of 3 to 8 metres with more than 15 subjects. Every thermal frame has a spatially registered visible-light counterpart sharing the same filename stem. CONTENTS thermal/ and rgb/ hold 1,895 registered PNG pairs each at 680x576. labels/ holds pose annotations as plain text, 17 lines of "x y" per person, plus a per-sample skeleton edge list. annotations_coco/ provides the same labels in COCO keypoint format, as the full set and as the fold-0 train (1,516 images) and validation (379 images) splits. splits/ contains the scenario-aware 5-fold manifests. inter_annotator_agreement/ contains independent second-annotator labels for a stratified 144-image subset. annotation_previews/ contains rendered skeleton overlays, which are derived and not required. code/ contains the annotation tools, the validation and benchmark harness computing MPJPE, PCK and OKS, the fold builder, the format converters, and code/analysis/ with the registration and fold-leakage measurement scripts. ACTIVITY CLASSES Calling, Drinking, Eating, Fighting, Interaction, Laptop, Phone Use, Reading, Smoking, Unconscious, Workout, and a Miscellaneous category for transitional or ambiguous poses. Class sizes are unbalanced, from 110 images for Reading to 240 for Miscellaneous. ANNOTATION QUALITY A stratified 144-image subset was independently re-annotated by a second annotator without access to the primary labels. Over the 177 people both annotators identified, agreement is 2.12 px mean keypoint deviation, 99.7% PCK@0.2 and 0.923 mean OKS; at the stricter PCK@0.05, a tolerance of about 12.6 px, they still agree on 96.4% of keypoints. The two annotators did not always agree on how many people a frame contains. Annotator A labelled 183 people and annotator B 189, of which 177 pair; person count agreed exactly in 133 of 144 images (92.4%), an instance-level agreement of 95.2%. The keypoint figures therefore describe localisation agreement on jointly identified people. Occluded joints are annotator-inferred positions, not observations. Every person carries all 17 keypoints and there is no visibility or occlusion flag. TPD therefore cannot be used as-is to train or evaluate occlusion-aware methods needing a visibility target, nor to compute visibility-weighted OKS or PCK. ACQUISITION AND CROSS-MODAL ALIGNMENT Captured with a synchronised bi-spectral Hikvision DS-2TD2628T-3/QA. The thermal channel is a 256x192 vanadium-oxide uncooled focal plane array (12 um pitch, NETD below 40 mK), paired with a higher-resolution CMOS optical channel. The released 680x576 resolution is a common working grid, not the thermal sensor's native resolution; it interpolates the 256x192 array by roughly 2.7x horizontally and 3x vertically. The thermal frames are 8-bit display-normalised images, not radiometric data: intensity is relative and per-frame and cannot be converted to absolute temperature or compared across frames. The raw 16-bit stream was not retained. Residual RGB-to-LWIR alignment was measured directly on the released pairs by phase cross-correlation of gradient magnitude over the subject region, using an estimator first validated by recovering known injected shifts to within 0.11 px. Over the 68% of frames with sufficient cross-modal edge structure the median residual is 2.34 px on the working grid, with a sub-pixel systematic component and no dependence on subject distance. This is the same order as the 2.12 px inter-annotator disagreement. The estimate models translation only. Both channels were recorded at 25 Hz, but TPD releases individually selected still frame pairs, not continuous video, so temporal and video-based analysis is out of scope. SPLITS The 5-fold protocol partitions each class into contiguous capture blocks so that near-duplicate frames from one take do not straddle a fold boundary. Measured against a random split of the same images at the same fold sizes, this reduces validation frames having a near-duplicate in training from 5.6% to 3.0%. The splits are temporally but not subject-disjoint: the release carries no subject identifiers, so the same person may appear in both partitions of a fold, and results are within-cohort accuracy that may be optimistic for unseen people. REFERENCE BENCHMARK Fifteen pose-estimation models spanning lightweight real-time detectors, heatmap, transformer, regression, coordinate-classification and one-stage families, together with the Sapiens foundation model, were evaluated under a single documented protocol. Best zero-shot transfer is 12.1 px MPJPE; in-domain fine-tuning reduces top-down architecture error to 7.9 px. NOT INCLUDED Subject identifiers, per-image capture distance and environment tags, capture-session identifiers, camera calibration, per-keypoint visibility flags, contiguous video, and raw radiometric temperature values. See README.md for the complete format specification.



