LMentry
收藏资源简介:
LMentry是一个专为大型语言模型设计的基准测试数据集,由特拉维夫大学创建。该数据集包含25个简单的语言任务,旨在快速评估模型的能力和鲁棒性。这些任务对人类来说非常简单,如包含特定单词的句子写作或识别列表中属于特定类别的单词。数据集创建过程中,每个任务都遵循了特定的指导原则,确保任务的简单性和可解释性。LMentry的应用领域主要集中在零样本评估,旨在解决大型语言模型在简单任务上的表现问题,以及对输入扰动的鲁棒性测试。
LMentry is a benchmark dataset dedicated to large language models, developed by Tel Aviv University. It consists of 25 simple language tasks designed to rapidly assess model capabilities and robustness. These tasks are highly straightforward for humans, such as composing sentences containing specific words or identifying words belonging to a designated category from a given list. During the dataset's creation, each task adheres to specific guiding principles to ensure its simplicity and interpretability. The primary application scenarios of LMentry focus on zero-shot evaluation, aiming to address the performance issues of large language models on simple tasks and conduct robustness tests against input perturbations.

- 1LMentry: A Language Model Benchmark of Elementary Language Tasks特拉维夫大学 · 2022年



