Assets 2026 - Beyond LLMs: Small Multimodal Models for Semantic Accessibility Validation of Alternative Texts
收藏资源简介:
Description This dataset supports research on semantic web accessibility evaluation using multimodal language models, with a particular focus on WCAG 2.1 Success Criterion 1.1.1 (Non-text Content) and Technique G94, which require meaningful alternative text for non-text content. The dataset was created from real-world websites and combines webpage metadata, contextual information, language model assessments, model-generated alternative text, and human evaluations. It was developed to support both the fine-tuning and evaluation of multimodal language-vision models for accessibility-related tasks. Data sources Images and their associated web context were collected from websites belonging to three application domains: News E-commerce Entertainment & Travel The selected websites include major international platforms such as BBC, CNN, Sky News, The Guardian, Amazon, Etsy, eBay, Nike, Decathlon, National Geographic, Wired, Gizmodo, Science Museum, Louvre Museum, and others. Dataset structure Each record corresponds to the evaluation of a single image by one participant. Consequently, the same image may appear multiple times, once for each evaluator. Each record contains: Webpage metadata Page URL Image URL Original alternative text (alt-text) HTML context surrounding the image Immediate textual context Nearby textual context Page title Page description Page keywords Page heading hierarchy GPT-4o model assessment The original alt-text was evaluated by a teacher language model, producing: Accessibility assessment score (1–5) Accessibility judgment (Success / Warning / Failure) Natural-language explanation of the assessment Improved alternative text Teacher model identifier Fine-tuned model assessment The same image was evaluated by the fine-tuned multimodal model, producing: Accessibility assessment score (1–5) Accessibility judgment (Success / Warning / Failure) Natural-language explanation Generated alternative text Model identifier Human evaluation For each participant, the dataset includes: Quality score assigned to the original alt-text Alternative text proposed by the participant Quality score assigned to the GPT-4o generated alt-text Quality score assigned to the fine-tuned Gemma3-4B generated alt-text Self-reported experience in digital accessibility Intended use The dataset is intended for research on: semantic accessibility evaluation; automatic assessment of alternative text quality; alternative text generation; multimodal language-vision models; fine-tuning and knowledge distillation; human-AI agreement in accessibility assessment; benchmarking accessibility-oriented language models. Dataset characteristics The dataset enables multiple research scenarios, including: comparison between original and generated alternative text; comparison between proprietary and open-weight models; analysis of agreement between human evaluators and AI models; investigation of the relationship between accessibility expertise and human judgments; evaluation of multimodal models on semantic accessibility tasks. Scope The dataset focuses exclusively on WCAG 2.1 Success Criterion 1.1.1 (Non-text Content). It is designed for semantic accessibility research and should not be interpreted as a comprehensive benchmark covering all WCAG success criteria.



