ServiceNow/PrivacyAlign
收藏资源简介:
PrivacyAlign是一个人类标注的偏好数据集,用于训练和评估隐私对齐的工具使用代理。每个数据行包含同一代理场景下两个不同模型生成的候选最终动作,以及人类偏好标签和每个响应的隐私标注(如泄露和遗漏)。场景是合成的,包括用户名、电子邮件、记忆和工具轨迹都是生成的,不包含真实用户数据。数据集分为训练集(1,150行)和测试集(200行),适用于训练奖励模型和大型语言模型代理,以使其更符合人类隐私规范,并评估代理性大型语言模型在最终工具调用中是否泄露敏感上下文或遗漏有用的非敏感细节。
PrivacyAlign is a human-annotated preference dataset for training and evaluating privacy-aligned tool-use agents. Each row pairs two candidate final actions from different models for the same agentic scenario, along with human preference labels and per-response privacy annotations (leaks and omissions). The scenarios are synthetic, with generated user names, emails, memories, and tool trajectories, and no real user data is included. The dataset is split into train (1,150 rows) and test (200 rows) sets, intended for training reward models and LLM agents to align with human privacy norms, and evaluating agentic LLMs on privacy leakage and omission in tool calls.




