sqlautophagycode_D_test_Qwen3-8B_strategy_trust_t1.0_g1_run0
收藏资源简介:
该数据集是一个包含1034个训练样本的资源,主要用于数据库查询、代码生成或问题解答相关任务。每个样本包含多个字段:任务ID(task_id)用于唯一标识任务;数据库ID(db_id)标识关联的数据库;提示词(prompt)提供任务背景或指令;问题(question)描述需要解决的具体问题;解决方案(solution)给出问题的一种解答;规范解决方案(canonical_solution)提供标准或参考解答;top_k进展(top_k_progression)可能记录模型生成过程中的中间结果或候选方案排序;完整模式(complete_schema)可能包含数据库模式或任务结构的完整描述。数据集适用于SQL查询生成、代码补全、问题求解模型训练与评估等场景,旨在支持自然语言处理模型在这些领域的开发与评估。
This dataset consists of 1034 training samples and is primarily used for tasks related to database querying, code generation, or question answering. Each sample includes multiple fields: task_id for uniquely identifying tasks; db_id for identifying associated databases; prompt for providing task background or instructions; question for describing specific problems to solve; solution for offering one answer to the problem; canonical_solution for providing a standard or reference answer; top_k_progression, which may record intermediate results or candidate solution rankings during model generation; and complete_schema, which may contain a complete description of the database schema or task structure. The dataset is suitable for scenarios such as SQL query generation, code completion, and training and evaluation of problem-solving models, aiming to support the development and assessment of natural language processing models in these areas.
- 数据集名称:sqlautophagycode_D_test_Qwen3-8B_strategy_trust_t1.0_g1_run0
- 数据集地址:https://huggingface.co/datasets/stefanocarrera/sqlautophagycode_D_test_Qwen3-8B_strategy_trust_t1.0_g1_run0
- 数据特征:
task_id:任务ID,类型为int64。db_id:数据库ID,类型为string。prompt:提示内容,类型为string。question:问题描述,类型为string。solution:解决方案,类型为string。canonical_solution:标准解决方案,类型为string。top_k_progression:Top-K进展信息,类型为string。complete_schema:完整模式信息,类型为string。
- 数据集划分:
- 仅包含训练集(
train):包含1034个样本,占用30860498字节。
- 仅包含训练集(
- 数据集大小:下载大小为5232215字节,总数据集大小为30860498字节。
- 配置:
- 配置名为
default,训练数据文件路径为data/train-*。
- 配置名为




