遇见数据集

SynLayers/Synlayers-Data

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为SynLayers Data,是一个用于图像到文本和文本到图像任务的训练数据集,语言为英语,规模在10万到100万样本之间。数据集包含SynLayers训练数据,其中边界框标注训练注释以synlayers_bbox.json文件提供。数据以Parquet分片格式存储,每个分片包含约5000个样本,存储在data/train-*.parquet路径下。每个样本代表一个分层图像,包含以下字段:base_image(背景/基础PNG字节)、whole_image(合成后的PNG字节)、layers(图层PNG字节及标题和边界框信息)和metadata(经过清理的原始metadata.json负载)。数据集可通过Hugging Face的datasets库加载使用。

Named SynLayers Data, this is a training dataset for image-to-text and text-to-image tasks, supporting English as its language, with a total of between 100,000 and 1,000,000 samples. The dataset includes SynLayers training data, and the training annotations with bounding box labels are provided in the synlayers_bbox.json file. The data is stored in Parquet shard format, where each shard contains approximately 5,000 samples and is stored under the path data/train-*.parquet. Each sample corresponds to a layered image, comprising the following fields: base_image (background/base PNG bytes), whole_image (synthesized PNG bytes), layers (layer PNG bytes along with title and bounding box information), and metadata (cleaned raw metadata.json payload). This dataset can be loaded and utilized via the Hugging Face datasets library.

提供机构:
SynLayers
二维码
社区交流群
二维码
科研交流群
商业服务