相关数据集
SEACrowd/wit
Wit数据集是一个基于维基百科的大型多模态多语言数据集,包含3760万条实体丰富的图像-文本对,涵盖108种维基百科语言。每种语言至少有1.2万个样本,其中53种语言有10万对图像-文本。该数据集特别关注东南亚地区的九种语言,包括ceb、fil、ind、jav、zlm、mya、tha、vie和war。数据集支持图像描述任务,并提供了多种加载方式,包括使用`datasets`库和`seacrowd`
Hugging Face2024-06-24 更新220
jp1924/KoreanVisionDataforImageDescriptionSentenceExtractionandGeneration
--- dataset_info: features: - name: id dtype: int32 - name: image dtype: image - name: caption dtype: string - name: caption_ls list: string - name: category dtype: str
Hugging Face2024-05-30 更新120
Ancient Chinese Data
Ancient Chinese Words Data Photos to the public, and it is open source. Chinese: 5218 wordsOracle: 1016 wordsXiaozhuan: 5104 wordsBronze: 1856 wordsTW: 5218 wordsChu Text: 1633 wordsQin Character:
DataCite Commons2021-04-18 更新60
CVasNLPExperiments/VQAv2_sample_validation_google_flan_t5_xxl_mode_D_PNP_GENERIC_Q_rices_ns_1000
--- dataset_info: features: - name: id dtype: int64 - name: question dtype: string - name: true_label sequence: string - name: prediction dtype: string splits: - name: fe
Hugging Face2023-05-26 更新130
vidore/syntheticDocQA_energy_test_captioning
--- dataset_info: features: - name: query dtype: string - name: image dtype: image - name: image_filename dtype: string - name: answer dtype: string - name: page dtype:
Hugging Face2024-06-06 更新100



