遇见数据集

Hairstyle, Facial Expression, and Clothing Dataset for Multi-Attribute Human Image Classification V2

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

This dataset was developed for a modular deep learning pipeline for multi-attribute human image classification, focusing on three human appearance attributes: hairstyle, facial expression, and clothing style. It supports independent classification of each attribute while allowing their predictions to be combined within a unified human appearance profile. This version (V2) substantially expands the dataset released in V1 to show greater data diversity and statistically rigorous evaluation. The dataset now comprises 2,314 hairstyle images (842 Curly: 715 train / 127 test; 814 Straight: 692 train / 122 test; 658 Wavy: 559 train / 99 test), up from 994 in V1; 6,335 facial-expression images (3,108 Neutral: 2,641 train / 467 test; 3,227 Smile: 2,743 train / 484 test), up from 2,579 in V1; and 2,969 clothing-style images (1,529 Casual: 1,299 train / 230 test; 1,440 Formal: 1,224 train / 216 test), up from 2,634 in V1. The hairstyle dataset combines the original publicly accessible Kaggle dataset used in V1 with an independent hairstyle dataset to increase visual and demographic diversity. The facial-expression dataset was expanded with additional unconstrained, "in-the-wild" images from an independent Kaggle facial-expression dataset, broadening the range of lighting, pose, and expression intensity beyond the original curated set. The clothing-style dataset was expanded using verified non-duplicate images from an independent Kaggle clothing dataset. All original V1 sources (Pexels images and voluntarily contributed photographs) remain part of the dataset. The evaluation protocol has also changed from V1. Each attribute-specific dataset is now divided into a fixed, stratified 15% held-out test set and an 85% training/validation pool (the train counts above). The training/validation pool is further split into 5 stratified cross-validation folds, with the same held-out test set used for final evaluation in every fold. This replaces the single fixed 80:10:10 split used in V1, providing a statistically robust, cross-validated evaluation protocol rather than a single-split point estimate. The released resource contains the processed images organized by attribute, class, and partition (train/test), corresponding to this updated protocol. During dataset preparation, considerable effort was made in image selection, manual screening, class labeling, deduplication, and organization. For the newly added images, a verification pass confirmed class labels matched the source datasets' own annotations and excluded any file with an invalid or corrupted image header. The dataset is publicly shared to support research reproducibility, independent evaluation, and further development of lightweight and modular approaches for human attribute recognition and multi-attribute image classification.

创建时间:
2026-09-03
二维码
社区交流群
二维码
科研交流群
商业服务