Cas-SEAT Dataset
收藏资源简介:
Cas-SEAT数据集是由浙江大学和新加坡国立大学的研究团队构建的,旨在增强高效多模态大语言模型(EMLLMs)的自评估能力。该数据集通过使用开源EMLLMs生成高质量的推理样本,并结合短提示进行少量数据标注,优化了数据效用和训练效率。数据集的内容包括图像、问题和答案,主要用于提升EMLLMs的推理和自评估能力。Cas-SEAT数据集的构建过程涉及将推理和自评估任务分解为两个独立的短提示任务,以减少长提示对EMLLMs的负担。该数据集的应用领域主要集中在多模态任务中,旨在解决EMLLMs在自评估和推理能力上的瓶颈问题。
Cas-SEAT dataset was constructed by the research teams from Zhejiang University and the National University of Singapore, aiming to enhance the self-evaluation capabilities of efficient multimodal large language models (EMLLMs). This dataset optimizes data utility and training efficiency by generating high-quality inference samples using open-source EMLLMs and conducting few-shot data annotation with short prompts. The dataset comprises images, questions and answers, and is primarily used to improve the inference and self-evaluation capabilities of EMLLMs. The construction process of Cas-SEAT dataset involves decomposing the inference and self-evaluation tasks into two independent short-prompt tasks, so as to reduce the burden of long prompts on EMLLMs. Its application scenarios mainly focus on multimodal tasks, aiming to address the bottleneck problems in the self-evaluation and inference capabilities of EMLLMs.

- 1Cascaded Self-Evaluation Augmented Training for Efficient Multimodal Large Language Models浙江大学, 新加坡国立大学 · 2025年



