CFLUE
收藏资源简介:
CFLUE是一个专为评估大型语言模型在金融领域中文理解能力而设计的基准数据集。该数据集由阿里巴巴集团和苏州大学计算机科学与技术学院共同创建,包含超过38,000个多选题和16,000多个测试实例,涵盖文本分类、机器翻译、关系提取、阅读理解和文本生成等多种NLP任务。CFLUE旨在通过这些任务全面评估模型的性能,特别是在金融知识评估和应用评估方面。数据集的创建过程涉及从公开渠道获取的模拟考试题目和专业人员标注的真实数据源,确保了数据的质量和多样性。CFLUE的应用领域主要集中在提升金融领域中文NLP任务的模型性能,解决现有数据集在规模和多样性上的限制。
CFLUE is a benchmark dataset specifically designed to evaluate the Chinese language understanding capabilities of large language models in the financial domain. Developed jointly by Alibaba Group and the School of Computer Science and Technology, Soochow University, the dataset contains over 38,000 multiple-choice questions and more than 16,000 test instances, covering a wide range of NLP tasks including text classification, machine translation, relation extraction, reading comprehension, and text generation. CFLUE aims to comprehensively evaluate model performance through these tasks, with a particular focus on financial knowledge assessment and application evaluation. The dataset was constructed using simulated exam questions sourced from public channels and real data annotated by professional personnel, ensuring the quality and diversity of the data. The main application scenarios of CFLUE are focused on improving the performance of models for Chinese NLP tasks in the financial domain, addressing the limitations of existing datasets in terms of scale and diversity.

- 1Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset阿里巴巴集团, 苏州大学计算机科学与技术学院 · 2024年



