WebUIBench
收藏资源简介:
WebUIBench是一个综合基准,用于评估多模态大语言模型在WebUI到代码转换中的表现。该数据集包含21K高质量的问题-答案对,来源于超过0.7K真实世界的网站,系统性地评估了四个关键领域:WebUI感知、HTML编程、WebUI-HTML理解和WebUI到代码转换。
WebUIBench is a comprehensive benchmark for evaluating the performance of multimodal large language models on WebUI-to-code translation tasks. It includes 21K high-quality question-answer pairs sourced from over 700 real-world websites, and systematically assesses four critical domains: WebUI perception, HTML programming, WebUI-HTML comprehension, and WebUI-to-code translation.
WebUIBench 数据集概述
基本信息
- 数据集名称: WebUIBench
- 发布日期: 2024年5月20日
- 许可证: MIT
- 数据集大小: 21K 高质量问答对
- 数据来源: 超过 0.7K 真实世界网站
数据集简介
WebUIBench 是一个系统设计的基准测试,用于评估多模态大型语言模型(MLLMs)在四个关键领域的能力:
- WebUI 感知
- HTML 编程
- WebUI-HTML 理解
- WebUI-to-Code
数据集特点
- 包含多维子能力的评估框架
- 专注于网页生成结果以外的多方面评估
- 基于软件工程原则设计
评估结果
- 评估了 29 个主流 MLLMs
- 包括 7 个闭源模型(如 GPT-4o、Gemini-1.5 Pro、Claude-3.5-Sonnet)
- 包括 22 个开源模型(如 InternVL2.5 系列、Qwen2-VL 系列)
获取方式
- Hugging Face 数据集: https://huggingface.co/datasets/Tele-AI-MAIL/WebUIBench
- GitHub 仓库: https://github.com/MAIL-Tele-AI/WebUIBench
引用
bibtex @article{xx, title={WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code}, author={xx}, journal={arXiv preprint arXiv:xx}, year={2025} }
致谢
- 感谢 VLMEvalKit 提供的工具和实现




