BIM-SSD-V2 Dataset (Malaysian Sign Language)
收藏资源简介:
This dataset was developed for Malaysian Sign Language (Bahasa Isyarat Malaysia, BIM) to support research in sign language recognition and translation. There is a first version of the dataset that was successfully published known as the BIM-SSD-V1 dataset. BIM-SSD-V1 was developed with images and video data collected from four BIM signers in both controlled and uncontrolled environments. It has covered 4,858 video samples recorded using smartphone cameras. That dataset consists of alphabets, numbers, words and sentences, as outlined in the Sign Language Module for Dataset Development.pdf. While a second version, BIM-SSD-V2 was developed through merging the RGB frames for alphabets, numbers, and words, producing a total of 3,143 RGB frame folders (each with 2D keypoints extracted using MediaPipe) and 4,900 glosses with their natural language translations. BIM-SSD-V2 has been split into 2,877 train, 146 validation and 120 testing sets for recognition purpose, while translation also has been prepared with 4,500 train, 200 validation and 200 testing set. This dataset provides the first standardized and multimodal resource for Malaysian Sign Language, supporting both continuous sign-to-gloss recognition and gloss-to-text translation research.
本数据集面向马来西亚手语(Bahasa Isyarat Malaysia,简称BIM)开发,旨在支撑手语识别与翻译领域的相关研究。其首个正式发布版本为BIM-SSD-V1数据集。BIM-SSD-V1数据集采集自四名BIM手语使用者在受控与非受控环境下拍摄的图像与视频数据,共计包含4858条由智能手机录制的视频样本。该数据集涵盖字母、数字、单词与句子内容,具体说明详见《Sign Language Module for Dataset Development.pdf》文档。 而第二版BIM-SSD-V2则通过合并字母、数字与单词的RGB帧构建而成,共计生成3143个RGB帧文件夹(每个文件夹均包含通过MediaPipe提取的二维关键点数据),以及4900条带有自然语言翻译的手语语汇(gloss)。针对识别任务,BIM-SSD-V2被划分为2877条训练样本、146条验证样本与120条测试样本;针对翻译任务,该数据集同样划分出4500条训练样本、200条验证样本与200条测试样本。本数据集为马来西亚手语提供了首个标准化多模态资源,可支撑连续手语-语汇识别以及语汇-文本翻译两类研究工作。



