astrology-dataset-clean
收藏资源简介:
该数据集是一个多模态数据集,包含图像和文本数据。每个样本由以下字段构成:source(数据来源,字符串类型)、page(页码,整数类型)、image(图像数据)、text(文本内容,字符串类型)、caption(图像标题,字符串类型)、category(类别标签,字符串类型)。数据集仅包含训练分割(train split),共有560个样本,总大小约为623 MB。基于字段结构,该数据集可能适用于图像描述生成、图文匹配、多模态分类或信息检索等任务,但具体任务定义和背景需参考额外文档或数据内容进一步确认。
This dataset is a multimodal dataset containing image and text data. Each sample consists of the following fields: source (data source, string type), page (page number, integer type), image (image data), text (text content, string type), caption (image caption, string type), category (category label, string type). The dataset only includes the train split, with a total of 560 samples and an approximate size of 623 MB. Based on the field structure, this dataset may be suitable for tasks such as image caption generation, image-text matching, multimodal classification, or information retrieval, but specific task definitions and background require further confirmation by referring to additional documentation or data content.
数据集名称
astrology-dataset-clean
数据集地址
https://huggingface.co/datasets/Phonsiri/astrology-dataset-clean
数据集描述
这是一个经过清洗的占星学数据集,包含图像和文本数据。
数据特征
- source:字符串类型,表示数据来源。
- page:整数类型,表示页面编号。
- image:图像类型,存储图像数据。
- text:字符串类型,存储文本内容。
- caption:字符串类型,存储图像对应的标题或描述。
- category:字符串类型,表示数据所属类别。
数据划分
- 训练集(train):包含672个样本,数据大小约为806.62 MB。
数据集大小
- 下载大小:约803.91 MB
- 总大小:约806.62 MB
配置
- 默认配置(default):训练集数据文件路径为
data/train-*。





