pandas-matplotlib-seaborn-synth
收藏资源简介:
该数据集是一个包含5000个训练样本的结构化文本数据集,主要用于代码生成或代码相关自然语言处理任务。数据集包含四个核心字段:task字段表示任务类型或类别;dataset_description字段提供数据集的描述信息;code字段包含编程代码示例;text字段包含与代码相关的文本内容。数据格式为结构化文本,总大小约11.1MB,下载大小约3.4MB。该数据集适用于代码生成、代码理解、文本到代码转换等任务的研究与开发。
This is a structured text dataset containing 5000 training samples, primarily designed for code generation or code-related natural language processing tasks. It includes four core fields: the "task" field indicates the task type or category; the "dataset_description" field provides descriptive information about the dataset; the "code" field contains programming code examples; and the "text" field holds text content related to the code. The dataset adopts a structured text format, with a total size of approximately 11.1 MB and a download size of around 3.4 MB. This dataset is suitable for research and development of tasks such as code generation, code understanding, and text-to-code translation.
数据集概述
- 数据集名称:Escobar/pandas-matplotlib-seaborn-synth
- 数据集地址:https://huggingface.co/datasets/Escobar/pandas-matplotlib-seaborn-synth
- 数据集配置:default
特征字段
| 字段名 | 数据类型 | 描述 |
|---|---|---|
| task | string | 任务类型 |
| dataset_description | string | 数据集描述 |
| code | string | 相关代码 |
| text | string | 文本内容 |
数据划分
| 数据集划分 | 样本数量 | 数据大小 |
|---|---|---|
| train | 5000 | 11,104,414 bytes |
其他信息
- 下载大小:3,393,345 bytes
- 数据文件路径:data/train-*




