MASKGROUPS-2M, MASKGROUPS-HQ
收藏资源简介:
本研究引入了一种新型的视觉语言任务,称为全模式指代表达式分割(ORES),该任务要求模型根据文本或文本加参考视觉实体的提示生成一组相关的掩码。为此,我们提出了一个名为“参考任何分割掩码组”(RAS)的框架,它通过掩码中心的大多模态模型增强分割掩码的语义理解,并生成满足用户指定指令的掩码组。为了训练和评估ORES模型,我们创建了数据集MASKGROUPS-2M和MASKGROUPS-HQ,这些数据集包含由文本和参考实体指定的多样化掩码组。通过广泛评估,我们证明了RAS在我们的新ORES任务以及经典的指代表达式分割(RES)和广义指代表达式分割(GRES)任务上的优越性能。
This study introduces a novel vision-language task named Full-Modal Referring Expression Segmentation (ORES), which mandates models to generate a set of relevant masks based on textual prompts or prompts integrating text and reference visual entities. To address this task, we propose a framework called Referring Any Segmentation Mask Groups (RAS), which employs a mask-centered multimodal large model to enhance the semantic understanding of segmentation masks and generate mask groups that comply with user-specified instructions. For training and evaluating ORES models, we construct two datasets: MASKGROUPS-2M and MASKGROUPS-HQ, which contain diverse mask groups specified by text and reference entities. Through extensive experimental evaluations, we demonstrate that RAS achieves superior performance on our newly proposed ORES task, as well as the classic Referring Expression Segmentation (RES) and Generalized Referring Expression Segmentation (GRES) tasks.

- 1Refer to Anything with Vision-Language Prompts伊利诺伊大学香槟分校和Adobe · 2025年



