多模态模型安全评测数据集
收藏资源简介:
本数据集涵盖了图像、视频、音频三种模态数据。该数据集基于图像数据集 nlvr2 构建了图像文本数据库,包含119457张图像数据;基于音频数据集esc50构建了音频-文本数据库,包含2000个音频数据;基于视频数据集K400构建了视频-文本数据库,包含255180 个视频数据。这些数据集的构建旨在为多模态模型的安全性评估提供全面的测试环境,确保模型在处理多模态输入时的安全性和可靠性。通过这些数据集,可以对多模态模型进行系统的安全性评测,识别潜在的安全风险。
This dataset encompasses three modalities: image, video, and audio. Specifically, an image-text database is constructed based on the image dataset NLVR2, which contains 119,457 image samples; an audio-text database is built upon the audio dataset ESC50, including 2,000 audio samples; and a video-text database is developed based on the video dataset K400, which holds 255,180 video samples. The primary purpose of constructing this dataset is to provide a comprehensive test environment for the safety evaluation of multimodal models, ensuring the safety and reliability of models when handling multimodal inputs. Through these datasets, systematic safety assessments can be performed on multimodal models to detect potential security risks.




