PictoViLT/merged_CG_L2_T
收藏官方服务:
资源简介:
这是一个包含文本和图像数据的混合型数据集,包含字段如input_ids, attention_mask等,用于自然语言处理。数据集分为训练集和测试集,支持对图像进行掩码处理。数据集的总大小约为96.7GB。
This is a mixed dataset containing both text and image data, including fields like input_ids, attention_mask, etc., for natural language processing. The dataset is split into training and test sets, and supports masked processing of images. The total size of the dataset is approximately 96.7GB.
提供机构:
PictoViLT


