Trotquonalize/ksl-pose-dictionary-poc
收藏资源简介:
KSL Pose Dictionary (PoC) 是一个用于韩国手语文本到姿态生成研究的关节点数据集,属于原型验证阶段。它包含来自两个主要来源的混合关节点数据:一是韩国国立国语院韩国手语词典的1,444个单词的关节点(使用OpenPose 137格式,通过RTMW-DW-L-M模型提取),二是NIASL2021数据集中灾难安全领域的2,287个基础手势片段的关节点(原始OpenPose 137格式)。数据集整合成一个包含4,511个独特手势的混合索引,并提供了用于第一阶段训练的语料库,包含20,085个样本(训练集18,067、验证集986、测试集1,032),每个样本由韩语句子和对应的KSL手势序列组成。此外,还包括一个包含2,984个唯一基础词元的KSL词汇表。该数据集主要用于支持基于字典的韩国手语生成原型开发,覆盖了NIASL2021原始手势间隔的93.3%。数据以JSON格式存储,包含身体、面部和手部的137个关节点坐标,适用于序列到序列模型训练和关节点可视化任务。
KSL Pose Dictionary (PoC) is a keypoint dataset for Korean sign language (KSL) text-to-pose prototype development. It consists of hybrid keypoint data from two main sources: 1,444 word keypoints from the National Institute of the Korean Languages Korean Sign Language Dictionary (in OpenPose 137 format extracted via RTMW-DW-L-M), and 2,287 base gloss segmentation keypoints from the NIASL2021 disaster safety domain (original OpenPose 137). The dataset integrates into a hybrid sign index with 4,511 unique signs and provides a Stage 1 training corpus of 20,085 samples (train 18,067 / val 986 / test 1,032), each containing a Korean text and corresponding KSL gloss sequence. It also includes a KSL vocabulary of 2,984 unique base lemmas. Designed for dictionary-based KSL production research, it covers 93.3% of NIASL2021 raw gloss intervals. Data is stored in JSON format with 137 keypoints per frame (body, face, hands), suitable for seq2seq model fine-tuning and pose visualization tasks.





