基于凯撒密码的污染抵抗基准
收藏资源简介:
本数据集基于凯撒密码,旨在创建一个污染抵抗基准,以评估大型语言模型(LLMs)的逻辑推理、算术推理和泛化能力。数据集包含不同偏移量的凯撒密码和两种类型的明文(自然语言英语单词和随机非单词),共计200条数据。凯撒密码的动态性质使得LLMs难以记住所有可能的查询,从而提高了评估的公平性。数据集轻量化,易于生成新实例,使得更新过程更为便捷。数据集可用于评估LLMs在不同任务上的能力,并揭示LLMs在污染控制下的真实性能。
This dataset is built upon the Caesar Cipher, designed to develop a pollution-resistance benchmark for evaluating the logical reasoning, arithmetic reasoning, and generalization capabilities of Large Language Models (LLMs). It comprises a total of 200 data entries, including Caesar Cipher texts with different shift offsets and two types of plaintexts: natural language English words and random non-words. The dynamic nature of the Caesar Cipher makes it challenging for LLMs to memorize all possible queries, thereby enhancing the fairness of the evaluation. The dataset is lightweight and facilitates the generation of new instances, making the update process more convenient. This dataset can be used to assess the performance of LLMs across various tasks and uncover the true performance of LLMs under contamination-controlled conditions.




