DataPajama
收藏资源简介:
DataPajama数据集是由浙江大学和阿里巴巴集团共同创建的,包含447亿个Token的预训练语料库。该数据集通过DataMan工具进行了质量评分和领域类型的标注,旨在优化大型语言模型的预训练过程。数据集涵盖了14个质量标准,包括准确性、连贯性、语言一致性、语义密度等,并分为15个常见应用领域,如医学、金融、法律等。DataPajama的构建是为了帮助大型语言模型在特定领域内提高上下文学习性能。
The DataPajama dataset was co-developed by Zhejiang University and Alibaba Group. It is a pre-training corpus containing 44.7 billion Tokens, and was scored for quality and annotated with domain categories via the DataMan tool, with the objective of optimizing the pre-training process of large language models. The dataset covers 14 quality metrics including accuracy, coherence, linguistic consistency, semantic density, and others, and is categorized into 15 common application domains such as medicine, finance, law, and more. The development of DataPajama aims to assist large language models in enhancing their in-context learning performance within specific domains.




