SamsungSDS-Research/Policy-on-the-Fly-Benchmark
收藏资源简介:
PoFBench(Policy-on-the-Fly Benchmark)是一个专为测试设计的基准数据集,旨在衡量基于策略的自定义过滤器在LLM(大型语言模型)驱动的系统中的性能。该数据集包含韩语(原始)和英语(翻译)的1,080个实例,覆盖8个策略类别和6种策略组合。数据集的设计核心是通过改变相同提示的策略组合来评估模型是否能够一致地改变其判定结果。PoFBench可用于评估LLM输入/输出过滤系统、AI代理护栏、聊天机器人内容过滤器以及RAG系统安全过滤器的性能。数据集包含有害内容,仅限用于AI安全研究目的。
PoFBench (Policy-on-the-Fly Benchmark) is a test-only benchmark designed to measure the performance of policy-based custom filters in LLM-powered systems. The dataset contains 1,080 instances — 540 in Korean (original) and 540 in English (translated), covering 8 policy categories and 6 policy combinations. The core design of PoFBench is to evaluate whether a model consistently changes its verdict when only the policy combination is changed for the same prompt. PoFBench can be used to evaluate the performance of policy-based filter systems operating in environments such as LLM input/output filtering systems, AI agent guardrails, chatbot content filters, and RAG system safety filters. The dataset includes harmful content and is intended solely for AI safety research purposes.




