NautData
收藏资源简介:
NautData是一个大规模的水下指令跟随数据集,包含145万对图像-文本,支持八种水下场景理解任务。数据集涵盖了图像、区域和对象级别的理解,拥有多样的会话结构和庞大的规模。它为水下场景理解模型的发展和评估提供了坚实的基础,有助于推动水下机器人探索、环境保护和资源开发等领域的发展。
NautData is a large-scale underwater instruction-following dataset containing 1.45 million image-text pairs, which supports eight underwater scene understanding tasks. It encompasses image-, region-, and object-level understanding, featuring diverse conversational structures and a substantial scale. This dataset provides a solid foundation for the development and evaluation of underwater scene understanding models, and facilitates the advancement of fields such as underwater robotic exploration, environmental protection and resource exploitation.
NAUTILUS 数据集概述
数据集基本信息
- 数据集名称: NautData
- 数据规模: 包含145万图像-文本对
- 数据类型: 大规模水下指令跟随数据集
- 主要用途: 支持水下大型多模态模型的开发和评估
数据集特点
- 专门针对水下场景理解任务构建
- 包含丰富的图像-文本配对数据
- 为水下视觉任务提供训练和评估基准
数据获取
- 数据地址: https://github.com/H-EmbodVis/NAUTILUS/tree/dataset
- 处理后的数据: https://huggingface.co/datasets/Wang017/NautData
- 标注文件: https://huggingface.co/datasets/Wang017/NautData-Instruct
相关模型
数据集支持以下两个版本的NAUTILUS模型:
- NAUTILUS(LLaVA): 基于LLaVA-1.5架构
- NAUTILUS(Qwen): 基于Qwen2.5-VL架构
模型性能
数据集在多个水下场景理解任务上表现出色,包括分类、描述、定位、检测、视觉问答和计数等任务。

- 1NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding华中科技大学 · 2025年



