遇见数据集

XL-WSD-LLM: Extending XL-WSD to evaluate Large Language Models

收藏
Zenodo2025-03-11 更新2026-05-26 收录
官方服务:

资源简介:

This benchmark extends XL-WSD. Starting from XL-WSD, we build a set of prompts for evaluating Large Language Models (LLMs) in two settings. The first is a multiple-choice task, and the second is a generative task in which we assess the quality of the generated definition. The benchmark consists of three compressed archives. Two archives contain training and test data for each task and language, while another is dedicated to the output of several LLMs that we evaluate. Each dataset includes data split into two folders: FT and TT. FT contains data without machine translation, while TT contains data where missing glosses are automatically translated. More details are available in the pre-print article "Exploring the Word Sense Disambiguation Capabilities of Large Language Models," published on arXiv.org.

本基准数据集拓展了XL-WSD。我们以XL-WSD为基础,构建了一套用于评估大语言模型(Large Language Models,LLMs)的提示词集,涵盖两类评估场景:其一为多项选择任务,其二为生成式任务,我们将在此类任务中评估模型生成的释义质量。 该基准数据集包含三个压缩归档文件。其中两份归档分别存储各任务与各语言对应的训练及测试数据,剩余一份归档用于存储本次评估所涉及的多款大语言模型的输出结果。每个数据集均被划分为FT与TT两个文件夹:FT文件夹内的数据未经过机器翻译,TT文件夹内的数据则对缺失的释义完成了自动翻译。 更多详细信息可参阅发布于arXiv.org的预印本论文《探索大语言模型的词义消歧能力》。

提供机构:
University of Bari
创建时间:
2025-03-11
二维码
社区交流群
二维码
科研交流群
商业服务