CAM3D
收藏资源简介:
Cam3D consists of 108 labelled videos of 12 mental states including spontaneous facial expressions and hand gestures. It was labelled using crowd-sourcing (inter-rater reliability Κ=0.45). We used three different sensors for data collection: Microsoft Kinect sensors, HD cameras, and microphones in the HD cameras. After the initial data collection, the videos were segmented. Each segment showed a single event such as a change in facial expression, head and body posture movement or hand gesture. From videos with public consent, a total of 451 segments were collected. The mean duration is 6 seconds. Labelling was based on context-free observer judgment. Public segments were labelled by community crowd-sourcing. Out of the 451 segmented videos we wanted to extract the ones that can reliably be described as belonging to one of the 24 emotion groups from the Baron-Cohen taxonomy. From the 2916 labels collected, 122 did not appear in the taxonomy so were not considered in the analysis. The remaining 2794 labels were grouped as belonging to one of the 24 groups plus agreement, disagreement, and neutral. To fi lter out non-emotional segments we chose only the videos that 60% or more of the raters agreed on. This resulted in 108 segments in total. The most common label given to a video segment was considered as the ground truth. The data is categorized by the ground-truth label and divided into seven folders. For each video segment, we provide the colour video, camera parameters, colour images and their corresponding aligned depth images.
Cam3D数据集包含108段标注视频,涵盖12种心理状态下的自发面部表情与手部动作。该数据集采用众包(crowd-sourcing)方式完成标注,其评分者间信度Kappa系数为Κ=0.45。数据采集阶段我们使用了三类传感器:微软Kinect(Microsoft Kinect)传感器、高清摄像头,以及内置在高清摄像头中的麦克风。初始数据采集完成后,我们对原始视频进行分段处理,每一段仅包含单一事件,例如面部表情变化、头部与躯干姿态移动或手部动作。从获得公开授权的视频中,我们共提取得到451个视频片段,其平均时长为6秒。标注工作基于无语境的观察者判断开展,公开授权的视频片段由社区众包完成标注。在这451个分段视频中,我们希望提取出可被可靠归类至巴伦-科恩情感分类体系(Baron-Cohen taxonomy)下24种情感类别的片段。在收集到的2916条标注结果中,有122条未出现在该分类体系中,因此未纳入后续分析。剩余的2794条标注结果被划分为24种情感类别,以及同意、反对与中立三类标签。为过滤掉非情感类片段,我们仅保留了获得60%及以上标注者一致认可的视频,最终共得到108个视频片段。我们将每个视频片段获得频次最高的标注标签作为其基准真值(ground truth)。数据集将按照基准真值标签进行分类,并划分为7个文件夹。针对每个视频片段,我们均提供了彩色视频、摄像头参数、彩色图像,以及与之对齐的深度图像。




