MFTCXplain
收藏资源简介:
MFTCXplain是一个多语言基准数据集,用于通过仇恨言论多跳解释来评估大型语言模型(LLMs)的道德推理能力。数据集包含来自葡萄牙语、意大利语、波斯语和英语的3,000条推文,带有二元仇恨言论标签、道德类别和文本跨度级理由。该数据集旨在解决当前评估基准的两个主要不足:缺乏解释道德分类的注释,限制了透明度和可解释性;以及对英语的过度关注,限制了跨不同文化背景下道德推理的评估。
MFTCXplain is a multilingual benchmark dataset for evaluating the moral reasoning capabilities of Large Language Models (LLMs) through multi-hop explanations of hate speech. This dataset contains 3,000 tweets from Portuguese, Italian, Persian and English, annotated with binary hate speech labels, moral categories, and text-span-level justifications. It aims to address two major shortcomings of current evaluation benchmarks: the lack of annotations that explain moral classifications, which limits transparency and interpretability; and the over-reliance on English, which restricts the evaluation of moral reasoning across different cultural contexts.

- 1MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Hate Speech Multi-hop ExplanationUniversity of Southern California, University of São Paulo, Saarland University, University of Melbourne, Howard University, Portland State University, Leiden University · 2025年



