open-phi/textbooks
收藏资源简介:
该数据集是一个利用大型语言模型(LLMs)创建的综合性开源资源库,类似于古代亚历山大图书馆。数据集包含多种来源的样本,包括使用RAG模型参考Wikipedia或其他搜索数据生成的样本、完全合成的样本以及使用GPT-3.5和GPT-4生成的样本。数据集的特征包括主题、模型、概念、大纲、Markdown格式、领域、子领域和RAG等信息。训练集包含1795个示例,总大小为397014633字节。
This dataset is a comprehensive open-source repository built with Large Language Models (LLMs), analogous to the ancient Library of Alexandria. It includes samples from diverse sources: those generated by RAG models referencing Wikipedia or other search data, fully synthetic samples, and samples produced using GPT-3.5 and GPT-4. The dataset's features include information such as topic, model, concept, outline, Markdown format, domain, sub-domain, and RAG. The training set comprises 1795 examples with a total size of 397,014,633 bytes.
数据集概述
本数据集旨在构建一个类似于历史上的亚历山大图书馆的全面开源知识库,利用大型语言模型(LLMs)来实现这一目标。



