MLissard
收藏资源简介:
MLissard数据集由国立圣保罗大学电气与计算机工程学院创建,是一个多语言的基准测试数据集,旨在评估语言模型在处理和生成不同长度文本的能力。该数据集包含300个示例,支持英语、德语、葡萄牙语、俄语、西班牙语和乌克兰语六种语言。数据集通过自动翻译系统和Python脚本生成合成数据,涵盖了对象计数、列表交集、最后一个字母连接和重复复制逻辑等任务。MLissard数据集的应用领域主要集中在自然语言处理中的长度泛化问题,旨在解决模型在处理长序列时性能下降的问题。
The MLissard dataset was created by the School of Electrical and Computer Engineering, University of São Paulo. It is a multilingual benchmark dataset designed to evaluate the capabilities of language models in processing and generating texts of varying lengths. This dataset contains 300 examples and supports six languages: English, German, Portuguese, Russian, Spanish, and Ukrainian. The synthetic data of this dataset is generated via automatic translation systems and Python scripts, covering tasks such as object counting, list intersection, last letter concatenation, and repetition and duplication logic. The MLissard dataset is primarily applied to the length generalization problem in natural language processing, aiming to address the performance degradation of models when processing long sequences.




