DROWZEE constructed benchmark dataset
收藏资源简介:
DROWZEE构建的基准数据集是一个利用逻辑推理和自动生成的时序逻辑规则来检测大型语言模型(LLM)中事实冲突性幻觉(FCH)的测试框架。该数据集通过爬取维基百科等知识库中的信息建立了一个全面的事实知识库,并自动生成问题答案对作为测试用例。数据集旨在解决LLM在处理涉及复杂逻辑关系和时序推理任务时的幻觉问题,适用于多个知识领域的LLM评估。
The benchmark dataset constructed by DROWZEE is a test framework for detecting factual conflicting hallucinations (FCH) in Large Language Models (LLMs) using logical reasoning and automatically generated temporal logic rules. This dataset builds a comprehensive factual knowledge base by crawling information from knowledge repositories such as Wikipedia, and automatically generates question-answer pairs as test cases. It aims to address the hallucination issues of LLMs when handling tasks involving complex logical relationships and temporal reasoning, and is suitable for LLM evaluation across multiple knowledge domains.




