Libra-Test
收藏资源简介:
Libra-Test是一个专为中文大模型护栏而构建的评测基准,涵盖了7个关键风险场景和超过5,700条专家标注的数据。数据集包括真实数据、合成数据和翻译数据,确保了数据的多样性和广泛性。真实数据来源于Safety-Prompts数据集,合成数据使用AART方法生成,翻译数据则来自BeaverTails测试集的翻译。数据集的使用步骤包括环境安装、数据加载、推理与评测等,提供了详细的脚本和参数说明。
Libra-Test is an evaluation benchmark specifically designed for Chinese large language model (LLM) safety guardrails. It covers 7 critical risk scenarios and over 5,700 expert-annotated data samples. The dataset comprises real-world data, synthetic data, and translated data to ensure the diversity and broad coverage of the corpus. Specifically, real-world data is sourced from the Safety-Prompts dataset, synthetic data is generated using the AART method, and translated data is derived from translations of the BeaverTails test set. The dataset's usage workflow includes environment setup, data loading, inference, and evaluation, with detailed scripts and parameter instructions provided.




