ship-dataset
收藏资源简介:
ShipBench是一个基于参数化生成的船舶结构图纸的元数据锚定视觉语言基准测试数据集,旨在评估视觉语言模型在船舶结构工程图纸上的推理能力。数据集包含6种商业船型(油轮、VLCC、散货船、集装箱船、LNG运输船、LPG运输船)和9个基于图纸的子任务(包括船型识别、扶强材类型、板厚、扶强材尺寸、货舱容积、命名板剖面面积、舱室定位、舱室边界、舱壁位置)。数据规模为6450个候选设计,每个设计包含剖面图和舱室图两个视角的PNG图像,总计12900张图像。核心基准测试包含5346个QA项(基于594个测试候选设计×9个子任务)。数据集在候选设计级别按照80/10/10的比例进行了分层划分(按船型,随机种子为42),具体划分为训练集5160个、验证集642个、测试集648个。所有任务的确定性真值均直接来源于生成器的输入字典和恢复的几何数据,无需人工标注。数据集还包含多个任务变体文件和预计算的模型预测结果及统计分析。
ShipBench is a metadata-anchored vision-language benchmark dataset based on parametrically generated ship structural drawings, designed to evaluate the reasoning capabilities of vision-language models (VLMs) on ship structural engineering drawings. The dataset covers 6 types of commercial ships: oil tanker, VLCC, bulk carrier, container ship, LNG carrier and LPG carrier, and includes 9 drawing-based subtasks: ship type identification, stiffener type, plate thickness, stiffener dimension, cargo hold volume, named panel section area, cabin positioning, cabin boundary and bulkhead position. The dataset comprises 6,450 candidate designs, each featuring two PNG images from different perspectives: section view and cabin view, with a total of 12,900 images. The core benchmark consists of 5,346 QA pairs, derived from 594 test candidate designs across the 9 subtasks. The candidate designs are stratified split at an 80/10/10 ratio by ship type, with a fixed random seed of 42, and are specifically divided into a training set of 5,160 samples, a validation set of 642 samples and a test set of 648 samples. The definitive ground truth for all tasks is directly sourced from the generator's input dictionary and recovered geometric data, eliminating the need for manual annotation. The dataset additionally provides multiple task variant files, pre-computed model prediction results and statistical analysis reports.





