first-impressions-v2
收藏资源简介:
First Impressions V2数据集包含了从超过3000个不同的YouTube高清视频中提取的10000个视频片段(平均时长15秒),这些视频中的面对镜头说英语的人展示了不同的性别、年龄、国籍和种族。视频被分为训练集、验证集和测试集,比例为3:1:1。视频通过亚马逊Mechanical Turk进行标注,标注包含五大性格特质变量,即五大人格模型(大五人格),包括外向性、宜人性、责任心、神经质和开放性。每个视频片段都有这五个特质的地面真实标签,值在0到1之间。数据集还扩展了新的语言数据(转录)和新的工作面试变量(面试注释)。转录由专业转录服务Rev完成,平均每个视频43个单词。此外,还有关于性别和种族的标注。数据集还提供了原始成对注释和预测属性(软标签)。
First Impressions V2 Dataset contains 10,000 video clips (average duration: 15 seconds) extracted from over 3,000 distinct high-definition YouTube videos, where English-speaking individuals facing the camera exhibit diverse genders, ages, nationalities and ethnicities. The dataset is split into training, validation and test sets with a ratio of 3:1:1. All videos were annotated via Amazon Mechanical Turk, with annotations covering the Big Five Personality Traits model, namely extraversion, agreeableness, conscientiousness, neuroticism and openness to experience. Each video clip has ground-truth labels for these five traits, with values ranging from 0 to 1. The dataset also includes new linguistic data (transcriptions) and novel job interview-related variables (interview annotations). The transcriptions were completed by professional transcription service Rev, with an average of 43 words per video. Additionally, annotations for gender and race are provided. The dataset also offers original pairwise annotations and predicted attributes (soft labels).




