Trust-videoLLMs
收藏资源简介:
Trust-videoLLMs数据集是一个用于评估视频大型语言模型(videoLLMs)可信度的综合基准,涵盖了真实、安全、鲁棒性、公平性和隐私五个维度。该数据集包含30个任务,涉及动态视觉场景、跨模态交互和现实世界的安全问题,旨在评估videoLLMs在多模态理解和分析方面的可信度。数据集由任务适配的现有数据集、使用高级文本/图像到视频工具生成的合成数据以及手动收集和标注的数据组成,确保了场景的多样性。该数据集为标准化可信度评估提供了一个公开可用的、可扩展的工具箱,填补了专注于准确性的基准和关键需求之间的差距。
The Trust-videoLLMs dataset is a comprehensive benchmark for evaluating the trustworthiness of video large language models (videoLLMs), covering five core dimensions: realism, safety, robustness, fairness, and privacy. It includes 30 tasks involving dynamic visual scenarios, cross-modal interactions, and real-world safety concerns, aiming to assess the trustworthiness of videoLLMs in multimodal understanding and analysis. The dataset is compiled from three sources: task-adapted existing datasets, synthetic data generated via advanced text/image-to-video generation tools, and manually collected and annotated data, which ensures the diversity of included scenarios. This dataset provides a publicly available and scalable toolkit for standardized trustworthiness evaluation, bridging the gap between accuracy-focused existing benchmarks and the critical unmet demand for rigorous trustworthiness assessment of videoLLMs.

- 1Understanding and Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding合肥工业大学, 清华大学 · 2025年



