PL-Guard
收藏资源简介:
PL-Guard是一个为波兰语语言模型安全分类而创建的手动标注基准数据集,旨在解决当前安全评估主要集中在高资源语言上的问题。该数据集包含了超过7,000个实例,主要由语言模型的回答组成,并经过专家评审进行安全标签标注。为了评估模型在不同语言环境下的鲁棒性,还创建了PL-Guard-adv,这是PL-Guard的对抗性扩展,具有文本扰动功能。PL-Guard-train包含6,487个实例,用于训练,而PL-Guard-test包含900个平衡测试实例。PL-Guard-adv-test是PL-Guard-test的扰动版本,用于评估模型在噪声输入下的鲁棒性。
PL-Guard is a manually annotated benchmark dataset developed for safety classification of Polish language models, aiming to address the gap that current safety evaluation studies primarily focus on high-resource languages. This dataset contains over 7,000 instances, which are mainly composed of responses from language models and annotated with safety labels through expert review. To evaluate the robustness of models across diverse linguistic contexts, PL-Guard-adv, an adversarial extension of PL-Guard with text perturbation functions, was also constructed. PL-Guard-train consists of 6,487 instances for model training, while PL-Guard-test includes 900 balanced test instances. PL-Guard-adv-test is a perturbed variant of PL-Guard-test, which is utilized to assess the robustness of models against noisy inputs.




