遇见数据集

Pattern Over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation

收藏
Zenodo2026-06-04 更新2026-05-26 收录
官方服务:

资源简介:

Multimodal large language models (MLLMs) are increasingly usedto translate webpage screenshots into front-end code, but repeatedUI patterns may sway them toward visually incorrect yet pattern-consistentoutputs. In this work, we test how repeated webpagepatterns hurt MLLM accuracy on an objective screenshot-to-codefill-in-the-blank task. We introduce the first benchmark for visualpattern-completion bias, where one localized element in a repeatedUI pattern is perturbed and the model must recover the maskedwidth or font-size value from the screenshot and HTML context.Starting from 30 webpages curated from Design2Code, we built1,440 evaluated screenshots spanning structural card and text-stylepatterns under standard and noise-overlaid conditions. We evaluatefive frontier MLLMs and find that all are strongly biased towardthe repeated baseline. Mean bias rate reaches 69.61% on card-widthperturbations and 80.22% on text font-size perturbations, whilemean accuracy is only 23.11% and 7.89%, respectively. Codex-5.3performs best but still drops from 65.83% accuracy on cards to13.89% on text, while Flash-3.0 reaches 96.11% bias on text. Noise,subtler perturbations, and boundary positions further increase biasrate. Reasoning analysis further shows that greater reasoning effortreduces bias, yet qualitative evidence reveals that models canidentify the anomalous element and still override it with the pattern-consistentanswer. Our results identify a concrete failure mode inmultimodal code generation and show that its severity is governedby visual saliency.

多模态大语言模型(Multimodal large language models, MLLMs)正日益被应用于将网页截图转换为前端代码,但重复UI模式可能会引导模型生成视觉上不准确却与模式一致的输出。本研究旨在探究重复网页模式如何损害MLLM在客观的截图转代码填空任务中的准确性。我们首次提出了针对视觉模式补全偏差的基准测试:在该测试中,重复UI模式中的某一局部元素会被扰动,模型需要从截图与HTML上下文中恢复被掩码的宽度或字体大小值。 我们从Design2Code中精选的30个网页出发,构建了涵盖结构卡片与文本样式模式、覆盖标准条件与叠加噪声条件的1440张评估用截图。我们对5款前沿MLLM进行了评估,发现所有模型均强烈偏向于重复的基准模式。在卡片宽度扰动任务中,平均偏差率达到69.61%,在文本字体大小扰动任务中则为80.22%,而对应的平均准确率仅分别为23.11%与7.89%。Codex-5.3表现最优,但在卡片任务上的准确率仍从65.83%降至文本任务的13.89%;而Flash-3.0在文本任务上的偏差率高达96.11%。噪声、更细微的扰动以及边界位置会进一步推高偏差率。 推理分析进一步表明,投入更多推理努力可降低偏差,但定性证据显示,模型即便能够识别异常元素,仍会以与模式一致的输出覆盖正确结果。我们的研究揭示了多模态代码生成中一种具体的失效模式,并表明其严重程度由视觉显著性所决定。

提供机构:
Zenodo
创建时间:
2026-03-31
二维码
社区交流群
二维码
科研交流群
商业服务