Content4All
收藏资源简介:
Content4All 是针对自动手语翻译研究的六个开放研究数据集的集合。手语翻译镜头由广播合作伙伴 SWISSTXT 和 VRT 拍摄。原始素材经过匿名处理,以提取 2D 和 3D 人体姿势信息。从大约 190 小时的处理数据中,发布了三个基本 (RAW) 数据集,即 1) SWISSTXT-RAW-NEWS、2) SWISSTXT-RAW-WEATHER 和 3) VRT-RAW。每个数据集都包含手语解释、相应的口语字幕和提取的 2D/3D 人体姿势信息。从每个基本数据集中选择一个子集并手动注释以对齐口语字幕和手语解释。子集选择类似于基准 Phoenix 2014T 数据集。我们的目标是让这三个新的带注释的公共数据集,即 4) SWISSTXT-NEWS、5) SWISSTXT-WEATHER 和 6) VRT-NEWS 成为基准并支持未来研究,因为该领域更接近于更大话语领域的翻译和生产.
Content4All is a collection of six open research datasets dedicated to automatic sign language translation research. The sign language translation footage was filmed by its broadcast partners SWISSTXT and VRT. Raw source materials were anonymized to extract 2D and 3D human pose information. Out of approximately 190 hours of processed data, three foundational (RAW) datasets have been released: 1) SWISSTXT-RAW-NEWS, 2) SWISSTXT-RAW-WEATHER, and 3) VRT-RAW. Each dataset contains sign language interpretations, corresponding spoken subtitles, and the extracted 2D/3D human pose information. A subset was selected from each foundational dataset and manually annotated to align the spoken subtitles and sign language interpretations. The subset selection follows the approach of the benchmark Phoenix 2014T dataset. Our goal is for these three new annotated public datasets — 4) SWISSTXT-NEWS, 5) SWISSTXT-WEATHER, and 6) VRT-NEWS — to serve as benchmarks and support future research, as the field is moving closer to translation and production in larger-scale discourse domains.




