TED4C-L
收藏资源简介:
TED4C-L是一个面向文化感知手势生成研究的大规模多模态数据集,由意大利理工学院和热那亚大学的研究团队构建。该数据集包含来自印度、意大利、土耳其和日本四个文化区域的764位演讲者,总时长106小时,共计659,454个五秒重叠样本,涵盖音频、运动姿态和文本转录等多维度信息。数据采集过程严格筛选了TED演讲视频,确保演讲者使用母语且姿态清晰,通过OpenPose提取9个上身关键点,并结合音频频谱特征与LaBSE文本嵌入进行对齐处理。该数据集旨在解决跨文化手势生成的泛化问题,支持语音同步、文化一致性建模等前沿人机交互研究。
TED4C-L is a large-scale multimodal dataset for culture-aware gesture generation research, constructed by research teams from the Istituto Italiano di Tecnologia (IIT) and the University of Genoa. It contains 764 speakers from four cultural regions: India, Italy, Turkey, and Japan, with a total duration of 106 hours and a total of 659,454 5-second overlapping samples, covering multi-dimensional information including audio, motion gestures, and text transcriptions. During data collection, TED talk videos were strictly screened to ensure that speakers used their native language and had clear gestures. Nine upper-body key points were extracted via OpenPose, and alignment processing was performed by combining audio spectrum features and LaBSE text embeddings. This dataset aims to address the generalization issue of cross-cultural gesture generation, and supports cutting-edge human-computer interaction research such as speech synchronization and cultural consistency modeling.




