OceanCorpus
收藏资源简介:
OceanCorpus 是一个专门针对海洋领域整理的维基百科实体知识数据集,旨在支持知识注入、预训练和监督微调(SFT)任务。该数据集为纯文本模态,语言为英语。完整数据集包含 113,626 个条目,当前预览版本提供了 1,000 个随机抽样的样本。数据结构包含两个字段:输入(由实体名称和实体类型构成的提示)和输出(实体描述)。由于数据量较大(超过 113k 条目和 580k+ PDF 文件),完整数据集托管在 Google Drive 上。
OceanCorpus is a Wikipedia-based entity knowledge dataset specifically curated for the marine domain, designed to support tasks including knowledge injection, pre-training, and supervised fine-tuning (SFT). This dataset is purely text-based and uses English as its language. The full dataset contains 113,626 entries, while the current preview version provides 1,000 randomly sampled instances. Its data structure consists of two fields: the "Input" field, which is a prompt constructed from the entity name and entity type, and the "Output" field, which is the entity description. Given its substantial scale—over 113,000 entries and more than 580,000 PDF files—the full dataset is hosted on Google Drive.
OceanCorpus 数据集概述
数据集简介
OceanCorpus 是一个专门为海洋领域构建的维基百科实体知识语料库。该数据集设计用于知识注入、预训练和监督微调。
关键特征
- 语言:英文。
- 模态:纯文本。
- 核心用途:知识注入、文本生成、预训练、监督微调。
- 领域:海洋。
数据结构与内容
- 特征:
input:字符串类型,由实体名称和实体类型构成的提示。output:字符串类型,实体描述。
- 数据划分:
train:包含 1000 个示例。
数据规模
- 总条目数:113,626(完整数据集)。
- 已上传样本数:1,000(用于预览的随机采样样本)。
- 完整数据规模:包含超过 113k 条条目和超过 580k 个 PDF 文件。
数据文件
data.csv:包含 1000 个采样条目。
完整数据集获取
由于数据集规模较大,完整语料库(CSV 文件及 PDF 文件)托管于 Google Drive。




