FELM
收藏资源简介:
FELM数据集是由香港科技大学开发的一个用于评估大型语言模型真实性的基准。该数据集收集了来自不同领域的响应,并进行了细致的真实性标注,旨在帮助研究人员和开发者识别和改进语言模型中的事实错误。数据集包含817个样本,覆盖了从世界知识到数学和推理等多个领域,通过细粒度的文本段落标注,可以精确地定位特定的事实错误。此外,数据集还提供了预定义的错误类型和参考链接,以支持或反驳声明,从而推动更可靠的语言模型的发展。
The FELM dataset is a benchmark developed by The Hong Kong University of Science and Technology for evaluating the factuality of large language models. It collects responses from diverse domains and conducts fine-grained factuality annotations, aiming to assist researchers and developers in identifying and rectifying factual errors in language models. The dataset contains 817 samples covering multiple domains ranging from world knowledge to mathematics and reasoning. Through fine-grained text paragraph annotations, it can accurately pinpoint specific factual errors. In addition, the dataset provides predefined error categories and reference links to support or refute claims, thereby advancing the development of more reliable language models.

- 1FELM: Benchmarking Factuality Evaluation of Large Language Models香港科技大学 · 2023年



