MuskumPillerum/General-Knowledge
收藏资源简介:
该数据集是一个围绕一般事实和推理主题的问题和答案集合。数据集分为两个特征 - 问题和答案。它旨在用于训练模型在一般知识和推理方面的能力。该数据集受到Alpaca数据集的启发,并且实际上包含了Alpaca数据集的一部分。数据分布详细列出了不同类别的问题和答案的比例,包括自然、人工智能、计算机科学、机器人技术、物理、化学、地理、历史、人物、体育等。数据集的语言为英语,使用MIT许可证。
This dataset is a collection of question-answer pairs focused on general facts and reasoning topics. It includes two core fields: question and answer. It is designed for training models to enhance their general knowledge and reasoning capabilities. Inspired by the Alpaca dataset, this collection actually contains a subset of the Alpaca dataset. The dataset's distribution details the proportional breakdown of question-answer pairs across various categories, including natural sciences, artificial intelligence, computer science, robotics, physics, chemistry, geography, history, notable individuals, sports, and more. The dataset is in English and is released under the MIT License.
数据集卡片 for General knowledge dataset
数据集概述
该数据集是一系列以常识和推理为主题的问题和答案集合。数据集分为两个特征:Question 和 Answer。旨在用于训练模型以擅长常识和推理。该数据集灵感来源于Alpaca数据集,实际上包含了Alpaca数据集的一个子集。
分布
数据集的总数(非Alpaca部分)为6315条,具体分布如下:
- 常识 - 80.8%
- 自然 - 16.5%
- 人工智能、计算机科学、机器人学 - 7.3%
- 物理、化学 - 16.3%
- 地理、历史 - 11.2%
- 人物 - 16%
- 体育 - 13.5%
- 推荐、推理、困境 - 17.8%
- 其他 - 1.4%
格式
数据集的格式示例如下: json { "Question": "What is the largest species of shark", "Answer": "The whale shark is considered the largest species of shark, with adults reaching lengths of up to 40 feet or more and weighing several tons." }
语言
英语
源数据
该数据集灵感来源于Stanford的Alpaca数据集:tatsu-lab/alpaca
许可信息
该数据集使用MIT许可证。
引用信息
目前,请引用:MuskumPillerum/General-Knowledge




