zjunlp/OceanCorpus
收藏资源简介:
OceanCorpus是一个大规模多模态数据集,旨在将结构化的海洋领域知识注入大型语言模型(LLMs)。它聚合了来自三个主要来源的数据:1) 网络知识(文本):从维基百科和权威海洋网站提取的113,626对问答数据;2) 论文知识(多模态):从约300篇同行评审的学术PDF中提取的高质量实体描述,包括图像路径和元数据;3) 开放数据集(图像):包含约44,810张特定领域图像,如珊瑚物种、野生鱼类和声纳目标。数据集支持文本生成、指令调整和视觉语言对齐。
OceanCorpus is a large-scale, multimodal dataset designed to inject structured marine domain knowledge into Large Language Models (LLMs). It aggregates data from three primary sources to support text generation, instruction tuning, and vision-language alignment: 1) Web Knowledge (Text-Only): A dataset of 113,626 instruction-style QA pairs extracted from Wikipedia and authoritative marine websites; 2) Paper Knowledge (Multimodal): High-quality entity descriptions extracted from approximately 300 peer-reviewed academic PDFs, including image paths and metadata; 3) Open-Dataset (Imagery): A collection of domain-specific images including coral species, wild fish, and sonar targets, totaling approximately 44,810 images.




