VC-RewardBench
收藏资源简介:
Visual-ERM 是一个用于视觉到代码任务的多模态生成奖励模型,直接在渲染的视觉空间中评估输出,并为结构化视觉重建提供细粒度、可解释且任务无关的差异反馈。该模型支持的任务包括图表到代码、表格到Markdown和SVG到代码。 关联数据集 VC-RewardBench 是一个用于评估结构化视觉数据上细粒度图像间差异判断的基准测试,包含1,335个精心策划的实例。每个实例包括一个真实图像、一个损坏/渲染的对应图像以及细粒度的差异标注。该数据集覆盖了图表、表格和SVG等多种视觉数据类型。 Visual-ERM 适用于视觉到代码的强化学习管道中的奖励建模、目标和预测渲染之间的视觉差异判断、推理时的基于反思的细化,以及视觉奖励建模和多模态强化学习的研究。
Visual-ERM is a multimodal generative reward model tailored for vision-to-code tasks. It evaluates model outputs directly in the rendered visual space, and delivers fine-grained, interpretable, task-agnostic discrepancy feedback for structured visual reconstruction. Supported tasks of the model include chart-to-code, table-to-Markdown, and SVG-to-code. The associated dataset VC-RewardBench is a benchmark for evaluating fine-grained inter-image discrepancy judgment on structured visual data, which contains 1,335 meticulously curated instances. Each instance comprises a ground-truth image, a corresponding corrupted/rendered image, and fine-grained discrepancy annotations. This dataset covers a variety of visual data types such as charts, tables, and SVGs. Visual-ERM can be applied to reward modeling in vision-to-code reinforcement learning pipelines, visual discrepancy judgment between ground-truth and predicted renderings, reflection-based refinement during inference, as well as research on visual reward modeling and multimodal reinforcement learning.



