synthetic-complaints-v2
收藏资源简介:
该数据集包含多个特征,如指令、输出、情感等,每个特征都有其特定的数据类型。数据集分为训练集和测试集,分别包含1812和460个样本。数据集的下载大小为295020字节,总大小为702486.746007891字节。数据集的配置名为'default',包含训练和测试数据文件。数据集的许可证为MIT,主要用于文本生成任务,语言为英语,数据集的友好名称为'Synthetic Complaints'。
This dataset encompasses multiple features including instruction, output, sentiment, and others, each with a specific data type. The dataset is split into training and test subsets, which contain 1812 and 460 samples respectively. The download size of the dataset is 295020 bytes, while the total size is 702486.746007891 bytes. The dataset's configuration is named 'default', which includes the training and test data files. The dataset is licensed under MIT, primarily intended for text generation tasks, uses English as its language, and has a friendly name of 'Synthetic Complaints'.
数据集概述
数据集名称
Synthetic Complaints
数据集信息
特征
- instruction: 字符串类型
- output: 字符串类型
- sentiment: 浮点数类型
- subjectivity: 浮点数类型
- word_count: 整数类型
- complexity_score: 浮点数类型
- style: 字符串类型
- topic: 字符串类型
数据分割
- train: 包含12392个样本,大小为3837107.1221398646字节
- test: 包含3140个样本,大小为971572.9667598633字节
数据集大小
- 下载大小: 1975645字节
- 数据集总大小: 4808680.088899728字节
配置
- config_name: default
- data_files:
- train: data/train-*
- test: data/test-*
- data_files:
许可证
MIT
任务类别
- 文本生成
语言
- 英语




