FaciaVox a Multimodal Biometric Dataset
收藏资源简介:
The FaciaVox dataset is an extensive multimodal biometric resource designed to enable in-depth exploration of face-image and voice recording research areas in both masked and unmasked scenarios. Features of the Dataset: 1. Multimodal Data: A total of 1,800 face images (JPG) and 6,000 audio recordings (WAV) were collected, enabling cross-domain analysis of visual and auditory biometrics. 2. Participants were categorized into four age groups for structured labeling:Label 1: Under 16 yearsLabel 2: 16 to less than 31 yearsLabel 3: 31 to less than 46 yearsLabel 4: 46 years and above 3. Sibling Data: Some participants are siblings, adding a challenging layer for speaker identification and facial recognition tasks due to genetic similarities in vocal and facial features. Sibling relationships are documented in the accompanying "FaciaVox List" data file. 4. Standardized Filenames: The dataset uses a consistent, intuitive naming convention for both facial images and voice recordings. Each filename includes:Type (F: Face Image, V: Voice Recording)Participant ID (e.g., sub001)Mask Type (e.g., a: unmasked, b: disposable mask, etc.)Zoom Level or Sentence ID (e.g., 1x, 3x, 5x for images or specific sentence identifier {01, 02, 03, ..., 10} for recordings) 5. Diverse Demographics: 19 different countries. 6. A challenging face recognition problem involving reflective mask shields and severe lighting conditions. 7. Each participant uttered 7 English statements and 3 Arabic statements, regardless of their native language. This adds a challenge for speaker identification. Research Applications FaciaVox is a versatile dataset supporting a wide range of research domains, including but not limited to:• Speaker Identification (SI) and Face Recognition (FR): Evaluating biometric systems under varying conditions.• Impact of Masks on Biometrics: Investigating how different facial coverings affect recognition performance.• Language Impact on SI: Exploring the effects of native and non-native speech on speaker identification.• Age and Gender Estimation: Inferring demographic information from voice and facial features.• Race and Ethnicity Matching: Studying biometrics across diverse populations.• Synthetic Voice and Deepfake Detection: Detecting cloned or generated speech.• Cross-Domain Biometric Fusion: Combining facial and vocal data for robust authentication.• Speech Intelligibility: Assessing how masks influence speech clarity.• Image Inpainting: Reconstructing occluded facial regions for improved recognition. Researchers can use the facial images and voice recordings independently or in combination to explore multimodal biometric systems. The standardized filenames and accompanying metadata make it easy to align visual and auditory data for cross-domain analyses. Sibling relationships and demographic labels add depth for tasks such as familial voice recognition, demographic profiling, and model bias evaluation.
FaciaVox数据集是一款大规模多模态生物特征资源,旨在支持戴口罩与未戴口罩场景下的人脸图像及语音录音相关研究的深度探索。 数据集特征: 1. 多模态数据:共采集1800张人脸图像(JPG格式)与6000条语音录音(WAV格式),可支撑视觉与听觉生物特征的跨域分析。 2. 年龄分组标注:将参与者划分为4个年龄组别并进行结构化标注,具体如下:标签1:16岁以下;标签2:16岁至30岁(不低于16岁且小于31岁);标签3:31岁至45岁(不低于31岁且小于46岁);标签4:46岁及以上。 3. 亲属样本集:部分参与者为亲属关系,由于其语音与面部特征存在遗传相似性,这为说话人识别(Speaker Identification, SI)与人脸识别(Face Recognition, FR)任务增加了挑战难度。亲属关系信息已在配套的“FaciaVox List”数据文件中记录。 4. 标准化文件名规范:数据集为人脸图像与语音录音采用统一且直观的命名规则,每个文件名包含以下信息:类型(F:人脸图像,V:语音录音)、参与者ID(例如sub001)、口罩类型(例如a:未戴口罩,b:一次性口罩等)、缩放级别或语句ID(例如图像采用1x、3x、5x,录音采用特定语句标识符{01, 02, 03, …, 10})。 5. 多样化人口统计学覆盖:数据集涵盖来自19个不同国家的参与者。 6. 高挑战性人脸识别场景:包含使用反光防护面罩遮挡以及极端光照条件下的人脸识别任务,具备较高研究挑战性。 7. 跨语言朗读任务:每位参与者需朗读7条英语语句与3条阿拉伯语语句,无论其母语为何种语言。这为SI任务增加了额外难度。 研究应用场景: FaciaVox是一款通用性极强的数据集,可支撑广泛的研究领域,包括但不限于: • 说话人识别(SI)与人脸识别(FR):在多变场景下评估生物特征识别系统的性能表现。 • 口罩对生物特征识别的影响:探究不同面部遮挡物对识别效果的作用机制。 • 语言对说话人识别的影响:探索母语与非母语语音对SI任务的影响。 • 年龄与性别估计:从语音与面部特征中推断人口统计学信息。 • 种族与族裔匹配研究:针对多样化人群开展生物特征相关研究。 • 合成语音与深度伪造检测:识别克隆或生成式语音。 • 跨域生物特征融合:融合面部与语音数据以实现更可靠的身份认证。 • 语音清晰度评估:量化口罩对语音清晰度的影响。 • 图像修复:重建被遮挡的面部区域以提升识别效果。 研究人员可独立或联合使用该数据集的人脸图像与语音录音,探索多模态生物特征识别系统。标准化的文件名与配套元数据便于将视觉与听觉数据进行对齐以开展跨域分析。亲属关系信息与人口统计学标签则为亲属语音识别、人口统计学画像构建以及模型偏见评估等任务提供了更丰富的研究维度。



