遇见数据集

BIM-SSD-V2 Dataset (Malaysian Sign Language)

收藏
Zenodo2025-12-17 更新2026-05-26 收录
官方服务:

资源简介:

This dataset was developed for Malaysian Sign Language (Bahasa Isyarat Malaysia, BIM) to support research in sign language recognition and translation. There is a first version of the dataset that was successfully published known as the BIM-SSD-V1 dataset. BIM-SSD-V1 was developed with images and video data collected from four BIM signers in both controlled and uncontrolled environments. It has covered 4,858 video samples recorded using smartphone cameras. That dataset consists of alphabets, numbers, words and sentences, as outlined in the Sign Language Module for Dataset Development.pdf. While a second version, BIM-SSD-V2 was developed through merging the RGB frames for alphabets, numbers, and words, producing a total of 3,143 RGB frame folders (each with 2D keypoints extracted using MediaPipe) and 4,900 glosses with their natural language translations. BIM-SSD-V2 has been split into 2,877 train, 146 validation and 120 testing sets for recognition purpose, while translation also has been prepared with 4,500 train, 200 validation and 200 testing set. This dataset provides the first standardized and multimodal resource for Malaysian Sign Language, supporting both continuous sign-to-gloss recognition and gloss-to-text translation research.

提供机构:
Zenodo
创建时间:
2025-11-13
二维码
社区交流群
二维码
科研交流群
商业服务