登录后查看消息通知
搜索
常见问题
消息
登录
首页
/
数据集
/
FuseCap-coco-karpathy-train-captions
FuseCap-coco-karpathy-train-captions
收藏
NIAID Data Ecosystem
2026-05-01 收录
图像描述生成
视觉语言
数据链接:
https://zenodo.org/record/8132275
数据链接
链接失效反馈
官方服务:
问题咨询
购买咨询
在线客服
NEW
资源简介:
Coco train captions produced by FuseCap.
应用场景:
创建时间:
2023-07-15
相关数据集
GEM
视觉语言
多模态数据
GEM是一个大规模的视觉-语言基准,包含用于图像-语言任务的GEM-I和用于视频-语言任务的GEM-V。它是目前覆盖图像-语言和视频-语言任务的最大视觉-语言数据集,并且标签化支持多种语言。
arXiv
2021-06-18 更新
16
0
VerboVision/VerboVision-Detail-Tags
细粒度图像标注
视觉语言
--- dataset_info: features: - name: image dtype: image: decode: false - name: instruction dtype: string - name: description dtype: string splits: - name: train
Hugging Face
2025-12-08 更新
8
0
lego_minifigure_captions
LEGO人物
图像描述生成
LEGO Minifigure Captions数据集包含12966张LEGO迷你人偶的图像及其描述。数据集包含以下列:'fig_num'表示迷你人偶的编号,'image'为JPEG格式的图像,'short_caption'为图像中迷你人偶的简短描述。数据来源于Rebrickable网站,图像从原始'minifigs.csv'文件的'img_url'列下载。未来计划添加使用Gemini-1.5-f
Hugging Face
2024-11-29 更新
6
0
5CD-AI/Vietnamese-BAAI-SVIT-llava-v1.5-format-gg-translated
多语言视觉问答
视觉语言
该数据集专注于视觉问答和问答任务,支持英语和越南语,数据规模在10万到100万之间,主要文件为SVIT_core_150K_vi.parquet。
Hugging Face
2024-06-14 更新
12
0
"MVVLP"
自动代客泊车
视觉语言
"Autonomous Valet Parking (AVP) represents a critical application of autonomous driving; however, existing approaches remain constrained by limited scene understanding, insufficient instruction compre
DataCite Commons
2025-10-03 更新
7
0
© 2023-2026 上海数据发展科技有限责任公司 版权所有
沪ICP备17003045号-15
沪公网安备31010402336585号
热门搜索
社区交流群
科研交流群
商业服务
数据资源
寻源服务
数据采集
标注服务
数据产品
代理销售
数据领域
凭证登记
数据产品
介绍推广