相关数据集
cjfcsjt/AITW_Single
--- dataset_info: - config_name: unseen_subject features: - name: ep_id dtype: string - name: step_id dtype: int64 - name: android_api_level dtype: int64 - name: current_activity
Hugging Face2024-04-24 更新510
5CD-AI/Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
该数据集是一个翻译后的数据集,包含178,509个caption条目、960,791个开放式问答条目和196,198个多项选择问答条目。数据集的视频来源自lmms-lab/LLaVA-Video-178K仓库。数据集的任务类别包括视觉问答和视频文本到文本转换,语言为英语和越南语,标签为视频和文本,大小类别为1M到10M之间。
Hugging Face2024-10-22 更新250
Nano1337/CLOVE-scores-negative-VQAnotCLIP
该数据集包含100个训练样本,每个样本包含uid、image、caption、CLIPscore、VQAscore、CLIPscore_percentile、VQAscore_percentile和percentile_difference等特征。其中,uid是字符串类型,image是图像类型,caption是字符串类型,CLIPscore、VQAscore、CLIPscore_percentil
Hugging Face2024-06-16 更新200
ranjaykrishna/visual_genome
Visual Genome是一个数据集和知识库,旨在将结构化的图像概念与语言连接起来。它包含108,077张图像,5.4百万区域描述,1.7百万视觉问答,3.8百万对象实例,2.8百万属性和2.3百万关系。该数据集主要用于图像到文本、对象检测和视觉问答等任务,所有注释均使用英语。
Hugging Face2023-06-29 更新470
JinghuiLuAstronaut/MTVQA_KR
--- dataset_info: features: - name: image dtype: image - name: question dtype: string - name: answer dtype: string - name: id dtype: int64 splits: - name: train num_b
Hugging Face2024-05-24 更新210



