BDSL47_Updated_2026
收藏资源简介:
Dataset Description BdSL47_Updated_2026 is an updated and reorganized version of the original BdSL47 dataset, which is publicly available through the original publication: https://doi.org/10.1016/j.dib.2023.109799. The original BdSL47 dataset contains images collected from 10 users/signers, organized across separate directories. To facilitate reproducible machine learning experiments and address potential signer overlap between dataset subsets, the present version reorganizes the images into a signer-independent train, validation, and test split. Each signer is assigned exclusively to one split; therefore, no signer appears in more than one of the three subsets. Signer-Independent Split The dataset was divided as follows: Training: User_03, User_04, User_06, User_07, User_08, User_09 Validation: User_05, User_10 Testing: User_01, User_02 This split ensures that the same signer is not present across the training, validation, and test sets. The number of samples per spllit as follows: Train samples: 28234 Validation samples: 9457 Test samples: 9341 Class Organization Each split contains 47 classes, represented by the folder names sign00–sign36 for Bangla alphabet-related signs and sign0–sign9 for numerical signs. Some inconsistencies in the original folder naming were also curated during the reorganization. The corresponding class labels are: Class ID Class Name sign00 অ/য় sign01 আ sign02 ই/ঈ sign03 উ/ঊ sign04 র/ঋ/ড়/ঢ় sign05 এ sign06 ঐ sign07 ও sign08 ঔ sign09 ক sign10 খ/ক্ষ sign11 গ sign12 ঘ sign13 ঙ sign14 চ sign15 ছ sign16 জ/য sign17 ঝ sign18 ঞ sign19 ট sign20 ঠ sign21 ড sign22 ঢ sign23 ণ/ন sign24 ত sign25 থ sign26 দ sign27 ধ sign28 প sign29 ফ sign30 ব/ভ sign31 ম sign32 ল sign33 শ/ষ/স sign34 হ sign35 ং sign36 ঁ sign0 ০ sign1 ১ sign2 ২ sign3 ৩ sign4 ৪ sign5 ৫ sign6 ৬ sign7 ৭ sign8 ৮ sign9 ৯ Landmark-Based Feature Representation In addition to the reorganized image dataset, hand landmark features were extracted from the images using MediaPipe Hands. For each image, 21 hand landmarks were extracted, with three coordinates (x, y, and z) for each landmark, resulting in 63 numerical features per image. The extracted features are provided as three CSV files: train.csv validation.csv test.csv The corresponding Class_ID is provided in the final column as the Label. Dataset Structure The main dataset is organized as follows: BdSL_Updated_2026/ ├── train/ ├── validation/ ├── test/ └── bdsl47_csv/ The train, validation, and test directories contain the corresponding image classes. The bdsl47_csv directory contains the extracted MediaPipe landmark features. For users who only require the landmark-based feature representation, a separate dataset version is provided: BdSL_Updated_2026_csv/ ├── train.csv ├── validation.csv └── test.csv Source and Code This dataset is derived from the original BdSL47 dataset. Users of this updated version should cite the original dataset publication: Original dataset DOI: https://doi.org/10.1016/j.dib.2023.109799 The scripts and supporting code used for dataset organization and landmark extraction are available in the following GitHub repository: https://github.com/Mehedi2154901019/Real-Time-Continuous-Bangla-Sign-Language-Recognition-Models-Pipelines-and-Experiments/tree/main/Static%20One%20Hand



