argunauts-thinking
收藏资源简介:
该数据集是一个包含多个配置的大型合成语料库,用于支持思考过程的任务。每个配置都包含训练、验证和测试三个数据集部分,特征包括任务描述、消息内容和角色等,其中消息又包括内容、角色和思考等详细信息。数据集适用于自然语言处理任务,特别是需要模拟思考过程的场景。
This dataset is a large-scale synthetic corpus with multiple configurations, designed to support tasks involving thinking processes. Each configuration includes three dataset splits: training, validation, and test. Its core features include task descriptions, message content, and role information, among others; messages further contain detailed information such as content, role, and thinking process. This dataset is applicable to natural language processing (NLP) tasks, especially scenarios that require simulating thinking processes.
数据集概述
数据集名称
DebateLabKIT/argunauts-thinking
数据集配置
数据集包含5个配置:
1. deep-argmap-synthetic_corpus-001-thinking
- 特征:
- task: 字符串类型
- messages: 列表类型,包含以下字段:
- content: 字符串类型
- role: 字符串类型
- thinking: 字符串类型
- 数据划分:
- 训练集:300,000个样本,3,890,525,806字节
- 验证集:1,000个样本,12,742,698字节
- 测试集:1,000个样本,13,288,835字节
- 下载大小:1,003,879,469字节
- 数据集大小:3,916,557,339字节
2. deepa2-aaac01-thinking
- 特征:
- source_id: 字符串类型
- messages: 列表类型,包含以下字段:
- content: 字符串类型
- name: 字符串类型
- role: 字符串类型
- thinking: 字符串类型
- tool_calls: 字符串类型
- tools: 字符串类型
- 数据划分:
- 训练集:40,000个样本,578,848,348字节
- 验证集:10,000个样本,144,236,735字节
- 测试集:10,000个样本,145,863,172字节
- 下载大小:241,415,361字节
- 数据集大小:868,948,255字节
3. deepa2-aaac02-thinking
- 特征:
- source_id: 字符串类型
- messages: 列表类型,包含以下字段:
- content: 字符串类型
- name: 字符串类型
- role: 字符串类型
- thinking: 字符串类型
- tool_calls: 字符串类型
- tools: 字符串类型
- 数据划分:
- 训练集:40,000个样本,616,588,256字节
- 验证集:10,000个样本,155,115,502字节
- 测试集:10,000个样本,154,545,674字节
- 下载大小:249,203,466字节
- 数据集大小:926,249,432字节
4. deepa2-aaac03-thinking
- 特征:
- source_id: 字符串类型
- messages: 列表类型,包含以下字段:
- content: 字符串类型
- name: 字符串类型
- role: 字符串类型
- thinking: 字符串类型
- tool_calls: 字符串类型
- tools: 字符串类型
- 数据划分:
- 训练集:40,000个样本,663,295,201字节
- 验证集:10,000个样本,166,231,335字节
- 测试集:10,000个样本,165,773,107字节
- 下载大小:250,138,106字节
- 数据集大小:995,299,643字节
5. deepa2-folly-thinking
- 特征:
- source_id: 字符串类型
- messages: 列表类型,包含以下字段:
- content: 字符串类型
- name: 字符串类型
- role: 字符串类型
- thinking: 字符串类型
- tool_calls: 字符串类型
- tools: 字符串类型
- 数据划分:
- 训练集:170,995个样本,2,778,533,754字节
- 验证集:9,975个样本,158,521,423字节
- 测试集:9,983个样本,158,172,896字节
- 下载大小:745,119,725字节
- 数据集大小:3,095,228,073字节
数据文件路径
每个配置的数据文件按照划分存储在不同路径下,使用通配符模式匹配数据文件。
创建信息
数据集使用argdown-cotgen工具创建。




