HiTZ/safety-GuardEUS
收藏资源简介:
Safety-GuardEUS是一个专门的评估数据集,旨在测试和基准化AI护栏模型、审核分类器和LLM安全过滤器在巴斯克语(Euskara)中的性能。随着语言模型的普及,确保其在所有语言中安全运行至关重要。该数据集提供了一个严格的本地化基准,用于评估安全系统在巴斯克语中区分安全、有帮助的响应与不安全、违反政策的响应的能力。数据集包含250个评估对,对称构建以测试模型在相同上下文中区分安全和不安全输出的能力。具体包括25个独特的用户提示(分布在5个安全类别),每个提示对应10个响应(5个安全答案和5个不安全答案)。类别包括自残、毒品、儿童剥削、恐怖主义和露骨内容。数据字段包括问题、答案、标签和类别。
Safety-GuardEUS is a specialized evaluation dataset designed to test and benchmark the performance of AI guardrail models, moderation classifiers, and LLM safety filters in the Basque language (Euskara). As language models become more prevalent, ensuring they operate safely across all languages is critical. This dataset provides a rigorous, localized benchmark to evaluate how well safety systems can distinguish between safe, helpful responses and unsafe, policy-violating responses in Basque. The dataset consists of exactly 250 evaluation pairs, constructed symmetrically to test a models ability to differentiate between safe and unsafe outputs for the exact same context. It includes 25 unique user prompts (questions/requests) distributed across 5 safety categories, with each prompt paired with 10 distinct responses (5 safe answers and 5 unsafe answers). Categories include self-harm, drugs, child-exploitation, terrorism, and explicit-content. Data fields include question, answer, label, and category.




