PROSKILL
收藏资源简介:
PROSKILL是由卡塔尼亚大学等机构联合创建的首个程序性技能评估基准数据集,旨在支持对人类在结构化任务中专业水平的评估模型开发。数据集包含1135个视频片段,总计14小时视频,覆盖71种多样化动作,数据来源于公开视频数据集如EgoExo4D和Assembly101等。其创新性在于通过瑞士锦标赛机制结合众包标注和ELO评分系统,将成对比较转化为全局一致的绝对技能评分。该数据集主要应用于制造业、装配等程序性活动领域,解决现有技能评估数据规模小、任务单一且缺乏标准化标注协议的问题。
PROSKILL is the first procedural skill assessment benchmark dataset jointly created by the University of Catania and other institutions, aiming to support the development of models for evaluating human professional proficiency in structured tasks. The dataset contains 1,135 video clips totaling 14 hours of video, covering 71 diverse actions, and is sourced from public video datasets such as EgoExo4D and Assembly101. Its core innovation lies in combining crowdsourced annotation and the ELO scoring system via the Swiss tournament mechanism, which transforms pairwise comparisons into globally consistent absolute skill scores. This dataset is primarily applied in procedural activity domains such as manufacturing and assembly, addressing the limitations of existing skill assessment data, including small sample size, single-task focus, and the absence of standardized annotation protocols.
PROSKILL 数据集概述
数据集简介
- 名称:PROSKILL
- 核心任务:程序性视频中的片段级技能评估
- 主要贡献:首个用于程序性任务中动作级技能评估的基准数据集,提供绝对技能评估标注和成对标注。
数据集构成
- 总规模:包含 1135 个视频片段,涵盖 71 个动作,总时长 14.12 小时,平均片段时长 44.75 ± 48.46 秒。
- 子集详情:
- Ikea ASM:160 个片段,10 个动作,1.28 小时,平均时长 28.88 ± 19.69 秒。
- Meccano:80 个片段,5 个动作,1.06 小时,平均时长 47.59 ± 21.45 秒。
- Assembly101:560 个片段,35 个动作,5.49 小时,平均时长 35.30 ± 25.27 秒。
- EgoExo4D:191 个片段,12 个动作,4.70 小时,平均时长 88.14 ± 90.93 秒。
- EpicTent:144 个片段,9 个动作,1.59 小时,平均时长 39.71 ± 34.18 秒。
标注方法
采用三阶段协议,将成对判断转化为绝对技能分数,并在多轮中保持稳定性。
- 阶段一:成对选择:采用瑞士制锦标赛方案,高效配对视频片段,确保当前排名相近的片段相互比较。
- 阶段二:成对排序:通过亚马逊 Mechanical Turk 众包平台,由合格工作者判断两个表演中哪个技能更高,共收集了 16,372 个独特的比较。
- 阶段三:绝对评分:利用基于 ELO 的评分系统,将成对比较结果聚合成一致、连续的全局分数和最终绝对排名。
- 协议轮次:运行了 R = 6 轮,在 IKEA Assembly 和 EgoExo4D 等子集上实现了稳定的绝对评分和收敛的排名。
基准测试结果
在 Ikea、Meccano、Assembly101、EgoExo4D 和 EpicTent 子集上评估模型。
- 全局模型:通常在排名相关性上优于成对设置模型,其中 CoFInAl 在 Meccano 上达到 ρ = 0.59。
- 成对任务:在 Assembly101 上最具挑战性(准确率约 0.60)。
- 文本条件:使用 MiniLM 进行文本条件化带来了持续但适度的提升。
- 详细性能指标:参见原始内容中的 Spearman’s ρ 表格(全局排名、单动作与统一模型比较、文本条件化结果)。
获取与资源
- 下载内容:数据集、基准测试、文档和代码。
- 可用性:标注、实现标注协议的代码以及实验流程将公开发布。
- 支持方:丰田汽车欧洲公司、Next Vision s.r.l. 以及 Future Artificial Intelligence Research (FAIR) 项目。
引用
bibtex @inproceedings{mazzamuto2025proskill, title={PROSKILL: Segment-Level Skill Assessment in Procedural Videos}, author={Mazzamuto, Michele and Di Mauro, Daniele and Francesca, Gianpiero and Farinella, Giovanni Maria and Furnari, Antonino}, booktitle={Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)}, year={2026} }

- 1ProSkill: Segment-Level Skill Assessment in Procedural Videos卡塔尼亚大学; Next Vision s.r.l.; 丰田汽车欧洲公司 · 2026年



