遇见数据集

FLUB

收藏
arXiv2024-02-17 更新2024-06-21 收录
数据链接:
官方服务:

资源简介:

FLUB是由清华大学创建的一个高质量数据集,专注于评估大型语言模型(LLMs)对谬误理解的能力。该数据集包含844个精心挑选的狡猾问题,这些问题在人类看来容易理解,但对模型来说极具挑战性。FLUB的数据来源于中国知名的在线论坛“弱智吧”,该论坛以其狡猾和不合理的发帖而闻名。数据集的创建过程包括数据清洗和标注,确保了数据的质量和适用性。FLUB的应用领域主要集中在推动LLMs对谬误的理解能力,从而提高它们处理复杂现实世界问题的能力。

FLUB is a high-quality dataset developed by Tsinghua University, focusing on evaluating the capacity of large language models (LLMs) to comprehend fallacies. The dataset contains 844 carefully selected tricky questions that appear easy for humans to understand but are extremely challenging for models. The data of FLUB is sourced from "Ruozhiba Bar", a well-known Chinese online forum famous for its tricky and illogical posts. The creation process of FLUB includes data cleaning and annotation, ensuring the quality and applicability of the dataset. The main application areas of FLUB center on promoting LLMs' understanding of fallacies, thereby improving their ability to handle complex real-world problems.

提供机构:
清华大学
创建时间:
2024-02-17
搜集汇总
数据集介绍
FLUB 数据集图片
背景与挑战
背景概述
FLUB是一个专门用于评估大型语言模型理解诡辩文本能力的基准数据集,包含诡辩类型分类、谬误解释和多项选择答案选择等任务。该数据集采用结构化JSON格式,每个样本包含诡辩文本、问题判断、诡辩类型、正确解释以及多项选择选项和答案,适用于测试模型对复杂逻辑谬误的识别和推理能力。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务