AWAREEVAL
收藏资源简介:
AWAREEVAL数据集由Lehigh University创建,旨在通过包含二元、多选和开放式问题来评估大型语言模型(LLMs)在五个意识维度上的表现:能力、使命、情感、文化和视角。该数据集通过多种问题类型全面了解LLMs的行为,特别关注LLMs在理解自身作为AI模型身份、识别其能力和使命以及展示社会智能方面的能力。AWAREEVAL的应用领域涉及AI对齐和安全性,强调了在可信和伦理发展中LLMs意识的重要性。
The AWAREEVAL dataset was created by Lehigh University. It is designed to evaluate the performance of Large Language Models (LLMs) across five awareness dimensions: competence, mission, emotion, culture, and perspective, using binary, multiple-choice, and open-ended questions. This dataset provides a comprehensive understanding of LLM behaviors through diverse question types, with a particular focus on LLMs' abilities to comprehend their own identity as AI models, recognize their inherent capabilities and missions, and demonstrate social intelligence. Applications of AWAREEVAL cover AI alignment and safety, emphasizing the significance of LLM awareness in the trustworthy and ethical development of AI.



