MELABenchv1
收藏资源简介:
MELABenchv1是一个用于评估大型语言模型在低资源语言马尔他语上的性能的基准数据集。它包含了11个判别性和生成性任务,旨在帮助研究人员评估和发展语言技术。数据集包含了55个公开可用的语言模型,包括不同大小的模型和不同的训练方法,如预训练和指令微调。通过对这些模型的评估,研究结果表明,在预训练和指令微调过程中接触马尔他语的模型在下游任务上表现更好。此外,该数据集还提供了几个相对较小的微调模型,这些模型在某些任务上的表现优于所有包含在该研究中的大型语言模型。
MELABenchv1 is a benchmark dataset for evaluating the performance of large language models (LLMs) on Maltese, a low-resource language. It comprises 11 discriminative and generative tasks, designed to assist researchers in evaluating and advancing language technologies. The dataset includes 55 publicly available language models, covering models of varying sizes and diverse training approaches including pre-training and instruction fine-tuning. Evaluations of these models reveal that models exposed to Maltese during their pre-training and instruction fine-tuning stages achieve superior performance on downstream tasks. Furthermore, the dataset also provides several relatively small fine-tuned models that outperform all large language models included in this study across certain tasks.

- 1MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP马耳他大学人工智能系 · 2025年



