遇见数据集
官方服务:

资源简介:

该数据集是一个大规模的多模态数据集,涵盖了三种模态:文本、图像和语音,并包含真实内容和机器生成内容的对齐实例。该数据集构建自COCO、Flickr8K和Places205数据集,并融入了五种最先进的生成模型所生产的机器生成内容。每个模态都拥有超过245,000个对齐的实例对。该数据集的主要任务是检测机器生成的内容。

This is a large-scale multimodal dataset covering three modalities: text, image, and speech, which contains aligned instances of both real content and machine-generated content. It is constructed from the COCO, Flickr8K, and Places205 datasets, and incorporates machine-generated content produced by five state-of-the-art generative models. Each modality includes over 245,000 aligned instance pairs. The primary task of this dataset is machine-generated content detection.

二维码
社区交流群
二维码
科研交流群
商业服务