noun-attributes2
收藏资源简介:
该数据集包含47,720个训练样本,总大小约为41.75GB。每个样本包含以下字段:名词(字符串类型)、属性(字符串类型)、提示词(字符串类型)、图像文件路径(字符串类型)、图像数据(图像类型)、GPT验证结果(布尔类型)、GPT判断结果(JSON字符串类型)、Qwen验证结果(字符串列表类型)以及GPT-Qwen联合验证结果(字符串列表类型)。数据集采用单训练集分割结构,未提供明确的背景说明或应用场景描述。
This dataset contains 47,720 training samples with a total size of approximately 41.75 GB. Each sample includes the following fields: noun (string type), attribute (string type), prompt (string type), image file path (string type), image data (image type), GPT verification result (boolean type), GPT judgment result (JSON string type), Qwen verification result (string list type), and GPT-Qwen joint verification result (string list type). The dataset adopts a single training set split structure, with no explicit background explanation or application scenario provided.
数据集概述
数据集基本信息
- 数据集名称: noun-attributes2
- 发布者: nirmalendu01
- 数据集地址: https://huggingface.co/datasets/nirmalendu01/noun-attributes2
数据集结构与内容
- 配置名称: default
- 数据文件:
- 训练集: data/train-*
- 数据特征:
noun: 名词 (字符串类型)attribute: 属性 (字符串类型)prompt: 提示词 (字符串类型)image_file: 图像文件路径 (字符串类型)image: 图像数据 (图像类型)gpt_verify: GPT验证结果 (布尔类型)gpt_judge_json: GPT判断结果 (JSON字符串类型)qwen_verified: Qwen验证结果 (字符串列表类型)gpt_verify_qwen: GPT对Qwen的验证结果 (字符串列表类型)
数据集规模
- 训练集样本数量: 48,000
- 训练集大小: 42,174,847,685 字节
- 数据集总大小: 42,174,847,685 字节
- 下载大小: 42,164,994,272 字节
数据格式与可用性
- 数据格式: 图像与文本混合数据集
- 数据分割: 仅包含训练集
- 访问方式: 通过Hugging Face数据集库下载



