遇见数据集

MFQ-30-LLM

收藏
arXiv2025-09-30 收录
官方服务:

资源简介:

该数据集是一个旨在评估大型语言模型道德推理能力的基准数据集,采用二元的“同意/不同意”格式,针对单个道德陈述进行评估。该数据集涵盖了包括LLaMA-2和GPT-4在内的多种大型语言模型,重点关注它们在理解和回应道德陈述方面的表现。这项任务被称为道德推理评估。

This dataset is a benchmark designed to evaluate the moral reasoning capabilities of large language models (LLMs). It adopts a binary "agree/disagree" format to assess individual moral statements. This dataset includes multiple LLMs such as LLaMA-2 and GPT-4, focusing on their performance in understanding and responding to moral statements. This task is termed moral reasoning evaluation.

提供机构:
AGI Research
二维码
社区交流群
二维码
科研交流群
商业服务