掌上视界:手机UI描述数据集
收藏资源简介:
针对小参数量视觉模型在图像理解深度不足、对手机端实际使用场景能力有限的问题,本文基于上万张原始图像精筛选构建了包含 1,672 张样本的数据集。以 Qwen-VL-235B 作为教师模型、Qwen-VL-2B-Instruct 作为学生模型开展知识蒸馏训练,面向手机端典型使用场景对模型进行专项优化,以提升端侧部署的视觉理解能力与实用性。
To address the issues that small-parameter visual models lack sufficient depth in image understanding and have limited performance in real-world mobile scenarios, this paper constructs a dataset containing 1,672 samples through meticulous screening of over 10,000 raw images. We employ Qwen-VL-235B as the teacher model and Qwen-VL-2B-Instruct as the student model to conduct knowledge distillation training, and perform targeted optimization of the model for typical mobile terminal usage scenarios to improve its visual understanding capability and practicality for on-device deployment.




