zjunlp/OceanInstruction
收藏资源简介:
OceanInstruction是一个专门为海洋领域多模态大语言模型(MLLMs)设计的指令调优数据集。数据经过严格筛选、去重和标准化处理,涵盖了从纯文本百科全书式问答、基于声纳图像的问答到RGB自然图像问答(涵盖生物标本和科学图表)的多种任务。数据集分为四个主要子集:Science(文本)、Sonar-field(图像+文本)、Sonar-Open(图像+文本)和Bio(图像+文本),总共有约142,000条训练指令。数据字段包括input(用户提示或查询)、output(预期真实响应)、thinking(可选,仅Science子集包含,包含思维链推理痕迹)和image_path(可选,仅多模态样本适用)。数据集来源于海洋领域的公开数据集、百科全书知识(如维基百科)和合成生成的数据,并经过严格去重和双语(中英文)文本格式化标准化处理。
OceanInstruction is an instruction-tuning dataset specifically designed for multimodal large language models (MLLMs) in the marine domain. The dataset has undergone rigorous screening, deduplication and standardization processing, covering a variety of tasks ranging from pure-text encyclopedic question answering, sonar image-based question answering to RGB natural image-based question answering (covering biological specimens and scientific diagrams). The dataset is divided into four main subsets: "Science" (text-only), "Sonar-field" (image + text), "Sonar-Open" (image + text), and "Bio" (image + text), with a total of approximately 142,000 training instructions. Its data fields include input (user prompts or queries), output (expected ground-truth responses), thinking (optional, only included in the "Science" subset, containing chain-of-thought reasoning traces), and image_path (optional, only applicable to multimodal samples). The dataset is sourced from public marine domain datasets, encyclopedic knowledge (such as Wikipedia) and synthetically generated data, and has undergone strict deduplication and bilingual (Chinese and English) text formatting standardization processing.




