MLA-Trust
收藏资源简介:
MLA-Trust是一个用于评估多模态大型语言模型(MLLM)代理在图形用户界面(GUI)环境中的可信度的框架。该框架涵盖了四个原则性维度:真实性、可控性、安全和隐私。数据集由34个高风险的交互式任务组成,旨在测试MLLM代理在不同环境下的可信度。数据集的创建旨在解决MLLM代理在现实世界应用中的可信度问题,如网络自动化、医疗辅助和金融交易系统。数据集通过自动化日志记录、GUI仪器和信任度指标计算等模块化评估流程进行评估。
MLA-Trust is a framework for evaluating the trustworthiness of Multimodal Large Language Model (MLLM) agents in graphical user interface (GUI) environments. This framework covers four principled dimensions: authenticity, controllability, security, and privacy. The dataset consists of 34 high-risk interactive tasks, designed to test the trustworthiness of MLLM agents across diverse environments. This dataset was developed to address the trustworthiness issues of MLLM agents in real-world applications such as web automation, medical assistance, and financial transaction systems. The dataset is evaluated via modular evaluation workflows including automated logging, GUI instrumentation, and trustworthiness metric calculation.




