遇见数据集

DeCoDELab/SynCAD-200k

收藏
Hugging Face2026-05-08 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个包含代码和图像对的多模态数据集,用于训练和评估相关任务。数据集包含206,500个训练样本,每个样本由三个主要特征组成:sample_id(字符串类型,唯一标识样本)、code(字符串类型,表示代码片段或代码描述)和image(图像类型,表示与代码相关的图像)。数据以训练集分割形式提供,总大小约为2.44 GB,下载大小约为1.93 GB。数据集旨在支持代码生成、图像理解或多模态学习等应用,但具体任务目标未在README中明确说明。

This dataset is a multimodal dataset containing code and image pairs, designed for training and evaluating related tasks. It includes 206,500 training examples, each consisting of three main features: sample_id (string type, uniquely identifying the sample), code (string type, representing code snippets or code descriptions), and image (image type, representing images associated with the code). The data is provided in a train split format, with a total size of approximately 2.44 GB and a download size of approximately 1.93 GB. The dataset aims to support applications such as code generation, image understanding, or multimodal learning, though the specific task objectives are not explicitly stated in the README.

提供机构:
DeCoDELab
二维码
社区交流群
二维码
科研交流群
商业服务