LlavaGuard数据集
收藏资源简介:
LlavaGuard数据集是由达姆施塔特工业大学等机构创建的高质量视觉数据集,专注于安全风险评估。该数据集包含3200张经过人工标注的图像,涵盖广泛的安全分类,用于调整视觉语言模型(VLM)以识别和评估上下文相关的安全风险。数据集的创建过程涉及从Socio-Moral Image Database(SMID)和网络爬虫中收集图像,并通过人工和合成方式进行标注。LlavaGuard数据集主要应用于AI模型的数据集管理和内容审核,旨在提高AI模型在处理视觉内容时的安全性和合规性。
The LlavaGuard dataset is a high-quality visual dataset developed by institutions including Technische Universität Darmstadt, with a core focus on safety risk assessment. It contains 3,200 manually annotated images spanning a wide spectrum of safety categories, and is designed for fine-tuning Vision-Language Models (VLMs) to identify and assess contextually relevant safety risks. The dataset construction process involves collecting images from the Socio-Moral Image Database (SMID) and web-crawled sources, with annotations generated via both manual and synthetic approaches. The LlavaGuard dataset is primarily utilized for dataset management and content auditing in AI models, aiming to enhance the safety and compliance of AI systems when processing visual content.

- 1LLavaGuard: VLM-based Safeguards for Vision Dataset Curation and Safety Assessment达姆施塔特工业大学 · 2024年



