Aloe Beta
收藏资源简介:
Aloe Beta数据集是巴塞罗那超级计算中心(BSC-CNS)和加泰罗尼亚理工大学(UPC)联合开发的医疗领域大型语言模型(LLM)训练数据集。该数据集包含1.2M条医学数据指令和420K条通过LLM生成的医学数据指令,旨在提高模型在医疗领域的专业知识和响应用户指令的能力。数据集由高质量的医学数据集和通过LLM生成的合成数据组成,旨在解决医疗领域LLM开发中的数据不足问题,并提高模型的安全性和可靠性。
The Aloe Beta dataset is a large language model (LLM) training dataset for the medical domain, jointly developed by the Barcelona Supercomputing Center (BSC-CNS) and the Polytechnic University of Catalonia (UPC). It contains 1.2 million medical data instruction samples and 420,000 LLM-generated medical data instruction samples, with the goal of enhancing the model's professional medical expertise and its capacity to respond to user instructions. The dataset is composed of high-quality medical datasets and LLM-generated synthetic data, designed to resolve the shortage of training data for medical LLMs and improve the safety and reliability of such models.

- 1The Aloe Family Recipe for Open and Specialized Healthcare LLMs巴塞罗那超级计算中心(BSC-CNS),西班牙 · 2025年



