Rapidata/svg-benchmark
收藏资源简介:
Rapidata静态SVG生成基准数据集由Rapidata构建,包含1,355,161条人类响应,用于比较30个前沿大型语言模型(LLM)从文本提示生成静态SVG(可缩放矢量图形)的能力。数据集包含188,754个头对头比较记录,每个记录对应两个模型对同一提示生成的SVG渲染结果,由人类标注者根据三个独立问题评分:偏好(主观视觉吸引力)、连贯性(视觉合理性,减少伪影)和对齐性(SVG与提示描述的匹配程度)。数据集涉及500个提示,生成14,872个独特的栅格化图像(PNG格式,分辨率768×768)。数据特性包括提示文本、两个图像、两个模型的SVG代码、模型ID以及三个评分维度的加权结果和详细结果。数据集的构建过程透明,包括提示集的多样性采样、30个模型的SVG生成、栅格化处理和人类评估。该数据集旨在评估LLM在SVG生成任务上的性能,并提供排行榜以支持模型比较和评估。
The Rapidata Static SVG Generation Benchmark dataset, built by Rapidata, contains 1,355,161 human responses comparing the performance of 30 frontier large language models (LLMs) in generating static SVGs (Scalable Vector Graphics) from text prompts. The dataset includes 188,754 head-to-head comparison rows, each representing two models SVG renders of the same prompt, scored by human annotators on three independent questions: Preference (subjective visual appeal), Coherence (visual soundness, fewer artifacts), and Alignment (faithfulness of the SVG to the prompt description). It involves 500 prompts, resulting in 14,872 unique rasterized images (PNG format, 768×768 resolution). Features include the prompt text, two images, SVG code from both models, model IDs, and weighted and detailed results for the three scoring dimensions. The dataset construction is transparent, covering diverse prompt sampling, SVG generation from 30 models, rasterization, and human evaluation. It is designed to evaluate LLM performance on SVG generation tasks and provides leaderboards for model comparison and assessment.




