Bench-CoE
收藏资源简介:
Bench-CoE数据集是由北京航空航天大学的人工智能研究所创建的,旨在评估和训练大型语言模型(LLMs)在多任务环境中的表现。该数据集包含多个领域的基准测试,如数学能力评估和视觉推理,用于训练路由器以分配任务给相应的专家模型。数据集的创建过程利用了现有的基准测试结果,通过查询级和主题级的标签来训练路由器。Bench-CoE数据集主要应用于多任务处理和跨领域推理,旨在提高模型在不同任务中的表现和泛化能力。
The Bench-CoE dataset was developed by the Institute of Artificial Intelligence at Beihang University to evaluate and train Large Language Models (LLMs) in multi-task scenarios. This dataset comprises benchmark tests across diverse domains, including mathematical proficiency assessment and visual reasoning, and is employed to train routers for assigning tasks to corresponding expert models. The development of this dataset leverages existing benchmark results, utilizing query-level and topic-level labels to train the routers. The Bench-CoE dataset is primarily utilized for multi-task processing and cross-domain reasoning, with the objective of enhancing model performance and generalization across various tasks.




