GMM-Sefai-Dataset
收藏资源简介:
该数据集是一个用于教育目的的图像标题数据集,源自大学项目。它包含164个训练样本,每个样本由图像及其相关文本描述组成。数据字段包括:图像文件、图像URL、由Qwen3-VL模型生成的初始标题、已验证的标题、模型标题以及数据划分标识。标题生成过程首先使用Qwen3-VL模型,后续可能经过验证或调整。数据集支持英语和立陶宛语两种语言,采用MIT许可证,总大小约为349MB。该数据集适用于图像标题生成、多语言自然语言处理等教育研究任务。
This dataset is an image captioning dataset for educational purposes, originating from a university project. It contains 164 training samples, each consisting of an image and its associated textual description. The data fields include: image file, image URL, initial caption generated by the Qwen3-VL model, verified caption, model caption, and data split identifier. The caption generation process first uses the Qwen3-VL model, with subsequent verification or adjustments possible. The dataset supports both English and Lithuanian languages, uses the MIT license, and has a total size of approximately 349MB. It is suitable for educational research tasks such as image caption generation and multilingual natural language processing.
- 数据集名称: GMM-Sefai-Dataset
- 许可证: MIT
- 语言: 英语 (en)、立陶宛语 (lt)
- 数据集大小: 约 349.88 MB(下载大小约 349.11 MB)
- 数据划分: 仅包含训练集(train),共 164 个样本
- 特征字段:
image: 图片数据(图像类型)image_url: 图片链接(字符串)qwen_caption: 使用 Qwen3-VL 生成的图像描述(字符串)verified_caption: 验证后的图像描述(字符串)model_caption: 模型生成的图像描述(字符串)split: 数据划分标识(字符串)
- 用途: 此数据集为大学项目的一部分,仅用于教育目的
- 联系: 如有任何投诉,可通过页面提及的方式联系




