head_qa_v2-aug
收藏资源简介:
该数据集是一个西班牙语文本数据集,包含两个核心文本字段:prompt(提示)和answer(答案)。数据集总大小约为14.73 MB,包含25,446个文本对样本,并划分为训练集(20,356个样本)和测试集(5,090个样本)。数据以字符串格式存储,适用于自然语言处理任务,如文本生成、问答系统构建、指令跟随模型训练或对话系统开发。其结构暗示了可能的应用场景:基于给定提示生成相应的答案文本。
This dataset is a Spanish text dataset containing two core text fields: prompt and answer. The total dataset size is approximately 14.73 MB, comprising 25,446 text pair samples, divided into a training set (20,356 samples) and a test set (5,090 samples). The data is stored in string format and is suitable for natural language processing tasks such as text generation, question-answering system construction, instruction-following model training, or dialogue system development. Its structure suggests potential application scenarios: generating corresponding answer text based on given prompts.
数据集概述
- 数据集名称:head_qa_v2-aug
- 配置名称:Spanish
- 数据集大小:约14.73 MB(下载大小约6.98 MB)
- 特征字段:
prompt(字符串类型)answer(字符串类型)
数据集划分
| 划分名称 | 样本数量 | 数据大小 |
|---|---|---|
| 训练集(train) | 20,356 条 | 11.81 MB |
| 测试集(test) | 5,090 条 | 2.92 MB |
数据文件路径
- 训练集:
Spanish/train-* - 测试集:
Spanish/test-*




