ljnlonoljpiljm/LVIS-extended
收藏资源简介:
该数据集是一个多模态数据集,包含图像及其相关元数据,适用于计算机视觉和自然语言处理任务。特征包括:图像ID、图像数据、高度、宽度、文件路径、数据划分(仅训练集)、图像标题(多个文本描述)、视觉问答(VQA)对(包含问题和答案),以及对象检测信息(如边界框、标签、分割掩码和同义词集)。数据集共有124801个训练示例,总大小约为7.7 GB,下载大小约为6.7 GB。
This dataset is a multimodal dataset containing images and associated metadata, suitable for computer vision and natural language processing tasks. Features include: image ID, image data, height, width, file path, data split (train only), image captions (multiple text descriptions), visual question answering (VQA) pairs (with questions and answers), and object detection information (such as bounding boxes, labels, segmentation masks, and synsets). The dataset consists of 124,801 training examples, with a total size of approximately 7.7 GB and a download size of approximately 6.7 GB.



