BLESS
收藏资源简介:
BLESS是一个评估大型语言模型在文本简化任务上性能的综合基准。该数据集由苏黎世大学的研究团队创建,包含44个不同大小、架构、预训练方法和可访问性的模型,针对三个不同领域的测试集(维基百科、新闻和医学)进行评估。BLESS旨在通过自动和手动分析,评估模型在少样本学习环境下的文本简化能力,特别是句子简化,以及模型执行的常见编辑操作的类型。该数据集的应用领域包括改进未来文本简化方法和评估指标的开发。
BLESS is a comprehensive benchmark for evaluating the performance of large language models (LLMs) on text simplification tasks. Developed by the research team from the University of Zurich, this benchmark encompasses 44 models with varying sizes, architectures, pre-training methodologies, and accessibility, and evaluates these models against three test sets spanning distinct domains: Wikipedia, news, and medicine. BLESS aims to assess the text simplification capabilities of models under few-shot learning settings, with a particular focus on sentence simplification, as well as the types of common editing operations performed by the models, through both automatic and manual analyses. The application scope of this benchmark includes advancing the development of future text simplification methods and their corresponding evaluation metrics.




