SciFaultyQA
收藏资源简介:
SciFaultyQA是一个专门用于评估大型语言模型(LLMs)识别和处理科学问题中错误能力的数据集。该数据集包含1333条科学问题,其中问题被故意设计为存在逻辑或科学上的错误。数据集的创建过程采用了GAN风格的合成数据生成方法,通过多个LLMs生成和验证错误问题。SciFaultyQA旨在解决LLMs在面对错误问题时无法识别其错误性质的问题,并为未来AI模型的基准测试提供新的方法。
SciFaultyQA is a dataset specifically developed to evaluate the capability of large language models (LLMs) to identify and handle errors within scientific questions. This dataset includes 1,333 scientific questions, each of which is intentionally crafted to contain logical or scientific errors. The dataset was constructed using a GAN-style synthetic data generation method, where multiple LLMs were utilized to generate and validate these erroneous questions. SciFaultyQA aims to address the issue that current LLMs fail to recognize the erroneous nature of flawed scientific questions, and provides a novel benchmarking approach for future AI model evaluations.
SciFaultyQA 数据集概述
数据集名称
SciFaultyQA
数据集描述
SciFaultyQA 是一个用于基准测试大型语言模型(LLMs)在检测错误科学问题能力的合成数据集。该数据集通过一种受生成对抗网络(GAN)启发的生成方法创建。




