CRiskEval
收藏资源简介:
CRiskEval是由天津大学创建的中文数据集,专门用于评估大型语言模型(LLMs)的风险倾向。该数据集包含14,888个问题,模拟了7种前沿风险类型,每个问题附带4个答案选项,均由人工标注风险级别。CRiskEval旨在通过细致的多项选择问答,测量LLMs在资源获取和恶意协调等方面的潜在风险。数据集的应用领域主要集中在评估和预防LLMs可能带来的风险,特别是在模型规模增大时,其对紧急自我维持和权力寻求等危险目标的倾向性增加。
CRiskEval is a Chinese dataset developed by Tianjin University, specifically designed to evaluate the risk propensity of large language models (LLMs). This dataset contains 14,888 questions simulating 7 cutting-edge risk categories, with each question paired with 4 answer options, and all options have their risk levels manually annotated. CRiskEval aims to measure the potential risks of LLMs in areas such as resource acquisition and malicious coordination via rigorous multiple-choice question-and-answer interactions. Its main application scenarios focus on assessing and mitigating the risks brought by LLMs, especially the elevated propensity for dangerous goals like emergent self-maintenance and power-seeking as model scales expand.




