UBIPose
收藏资源简介:
The UBIPose dataset is intended for the evaluation of head pose estimation algorithms in natural and challenging scenarios. This dataset provides the annotation of the positions of 6 facial landmarks (two corners of two eyes, nasal root and nose tip) in 14.4 K frames and 3D head poses (roll, pitch, yaw) in 10.4 K frames. Description of the Corpus The UBIPose dataset relies on videos from the UBImpressed dataset, which has been captured to study the performance of students from the hospitality industry at their workplace. The role play happens at a reception desk, where students interact with a research assistant who plays the role of a customer. Students and clients are recorded using a Kinect 2 sensor (one per person). In this free and natural setting, large head poses and sudden head motions are frequent as people are observed from a relatively large distance, and people are mainly seen from the side. Idiap Research Institute shares this dataset to enable the evaluation of head pose estimation algorithms in free and challenging scenarios. Out of the 160 interactions recorded in the UBImpressed dataset, we selected 32 videos. These videos are divided as follows: 22 videos (with 22 different persons) are provided as evaluation data. In 10 of these videos, 30-50 second clips were cut from the original videos and all frames were annotated. The other 12 videos were fully annotated at one frame persecond. This allowed to gather a large diversity of situations. In total, this amounts to 14.4K frames. The labels we provide are the positions of 6 facial landmarks (two corners of two eyes, nasal root and nose tip) and 3D head poses (roll, pitch, yaw). 10 additional videos can be used for processing and illustrating algorithmic results. These videos are unannotated and intended for the visualization of methods in scientific dissemination activities. Dataset content The dataset contains both the orignal video files to be processed (depth and RGB), the ground truth files (including those used for reconstruction and exploited for landmark localization evaluations), and code to evaluate performance. More precisely, the list is as follows: the RGB videos of the 22 test videos from 22 different users used in the paper for performance evaluation; the synchronized depth videos of these 22 test videos; the audio frame indices of these 22 test videos; the annotated landmarks for 14.4K frames; the validated inferred head poses for 10.4K frames; the full output results of our method; software code to allow computing the performance reported in the paper, as well as performance from produced pose results. videos for display: 10 additional pairs of RGB and synchronized depth videos can be used for processing and illustrating the algorithm results. These videos are unannotated and only intended for the visualization of methods in public dissemination activities. References @inproceedings{Muralidhar:2016:TJB:2993148.2993191, author = {Muralidhar, Skanda and Nguyen, Laurent Son and Frauendorfer, Denise and Odobez, Jean-Marc and Schmid Mast, Marianne and Gatica-Perez, Daniel}, title = {Training on the Job: Behavioral Analysis of Job Interviews in Hospitality}, booktitle = {Proceedings of the 18th ACM International Conference on Multimodal Interaction}, series = {ICMI 2016}, year = {2016}, location = {Tokyo, Japan}, pages = {84--91}, numpages = {8}, publisher = {ACM}, address = {New York, NY, USA} } @inproceedings{Yu:PAMI:2018, author = {Yu, Yu and Kenneth Alberto and Funes Mora and Odobez, Jean-Marc}, title = {HeadFusion: 360 Head Pose tracking combining 3D Morphable Model and 3D Reconstruction}, booktitle = {IEEE Transaction on Pattern Analysis and Maschine Intelligence (PAMI)}, year = {2018} }
UBIPose数据集(UBIPose Dataset)旨在评估自然且复杂场景下的头部姿态估计算法(head pose estimation algorithms)。该数据集为1.44万帧图像标注了6个面部关键点(双眼外眼角、鼻根与鼻尖)的位置,并为1.04万帧图像提供了三维头部姿态(滚转roll、俯仰pitch、偏航yaw)标注。 数据集概述 UBIPose数据集依托UBImpressed数据集(UBImpressed Dataset)的视频素材构建,后者采集自酒店服务业学生的工作场景,用于研究其在岗表现。实验环节为接待台角色扮演:学生与扮演顾客的研究助理进行互动,并使用Kinect 2传感器(每人一台)记录双方影像。在这种自由自然的真实场景中,由于拍摄距离相对较远且多为侧面取景,大角度头部姿态与突发头部动作十分常见。 IDIAP研究院(IDIAP Research Institute)公开该数据集,以支持自由复杂场景下头部姿态估计算法的评估工作。从UBImpressed数据集录制的160段交互素材中,我们筛选出32段视频,其划分方式如下:22段视频(对应22名不同受试者)作为评估数据集。其中10段视频从原始素材中截取了30至50秒的片段,并对所有帧进行标注;剩余12段视频则以每秒一帧的频率完成全量标注。此举确保了场景样本的丰富多样性,最终累计得到1.44万帧数据。我们提供的标注信息包含6个面部关键点坐标,以及1.04万帧的三维头部姿态(滚转、俯仰、偏航)数据。另有10段未标注视频可用于算法结果的处理与演示,供学术传播活动中可视化研究方法使用。 数据集内容 本数据集包含待处理的原始视频文件(深度图与RGB影像)、真值标注文件(含用于三维重建以及关键点定位评估的标注),以及性能评估代码。具体内容如下: 1. 用于论文中性能评估的22名受试者的22段测试视频的RGB影像; 2. 上述22段测试视频的同步深度视频; 3. 上述22段测试视频的音频帧索引; 4. 1.44万帧的标注面部关键点数据; 5. 1.04万帧的经验证的头部姿态推演结果; 6. 我们所提方法的完整输出结果; 7. 用于复现论文中报道的性能指标,以及计算自定义生成姿态结果性能的软件代码。 展示用视频 另有10段额外的RGB与同步深度视频对,可用于算法结果的处理与演示,此类视频未做标注,仅用于学术传播活动中的研究方法可视化。 参考文献 @inproceedings{Muralidhar:2016:TJB:2993148.2993191, author = {Muralidhar, Skanda and Nguyen, Laurent Son and Frauendorfer, Denise and Odobez, Jean-Marc and Schmid Mast, Marianne and Gatica-Perez, Daniel}, title = {在岗培训:酒店业面试行为分析}, booktitle = {Proceedings of the 18th ACM International Conference on Multimodal Interaction}, series = {ICMI 2016}, year = {2016}, location = {Tokyo, Japan}, pages = {84--91}, numpages = {8}, publisher = {ACM}, address = {New York, NY, USA} } @inproceedings{Yu:PAMI:2018, author = {Yu, Yu and Kenneth Alberto and Funes Mora and Odobez, Jean-Marc}, title = {HeadFusion:结合三维可变形模型与三维重建的360°头部姿态追踪}, booktitle = {IEEE Transaction on Pattern Analysis and Maschine Intelligence (PAMI)}, year = {2018} }




