R-Judge
收藏资源简介:
R-Judge是一个专为评估大型语言模型(LLMs)在处理代理交互记录时判断和识别安全风险能力而设计的数据集。该数据集由上海交通大学电子信息与电气工程学院创建,包含162条多轮代理交互记录,覆盖27个关键风险场景,涉及7个应用类别和10种风险类型。数据集中的每条记录都包含了用户指令和代理行动及环境反馈的历史,旨在通过这些复杂的多轮交互来评估LLMs的安全风险意识。R-Judge的应用领域主要集中在提高LLMs在开放代理场景中的安全风险意识,以解决在交互环境中可能出现的意外安全问题。
R-Judge is a specialized dataset designed to evaluate the ability of Large Language Models (LLMs) to judge and identify security risks when processing agent interaction logs. This dataset was created by the School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, and includes 162 multi-turn agent interaction records covering 27 critical risk scenarios, involving 7 application categories and 10 risk types. Each record in the dataset contains the historical sequence of user instructions, agent actions and environmental feedback, aiming to assess the security risk awareness of LLMs through these complex multi-turn interactions. The primary application of R-Judge is to enhance the security risk awareness of LLMs in open agent scenarios, thereby addressing potential unexpected security issues that may arise in interactive environments.

- 1R-Judge: Benchmarking Safety Risk Awareness for LLM Agents上海交通大学电子信息与电气工程学院 · 2024年



