nz_research_commons_gpt_oss_cross_provider_failed_rows_1200
收藏资源简介:
该数据集是一个结构化的学术文献数据集,包含24个训练样本。数据集中每条记录代表一篇学术论文,核心字段包括论文标题、作者、主题、摘要、全文文本以及唯一的记录哈希ID。数据集还包含丰富的标注信息,如分类标签、分类原因、分类置信度以及论文发表年份。为支持自然语言处理任务,数据集提供了提示词、对话内容、用于嵌入的文本以及对应的token数量。此外,数据集包含了基于GPT模型的自动化处理结果,包括模型输出、置信度、推理过程、错误信息及处理耗时,但部分GPT相关字段值为空。该数据集适用于多种NLP任务,如文本分类、对话生成、文本嵌入、学术文献分析以及大语言模型输出评估。
This dataset is a structured academic literature dataset containing 24 training samples. Each record in the dataset represents an academic paper, with core fields including paper title, authors, subject, abstract, full text, and a unique record hash ID. The dataset also includes rich annotation information, such as classification labels, classification reasons, classification confidence, and the year of publication. To support natural language processing tasks, the dataset provides prompts, dialogue content, text for embedding, and corresponding token counts. Additionally, the dataset contains automated processing results based on GPT models, including model outputs, confidence, reasoning processes, error messages, and processing time, although some GPT-related fields are empty. This dataset is suitable for various NLP tasks, such as text classification, dialogue generation, text embedding, academic literature analysis, and large language model output evaluation.
数据集概述:nz_research_commons_gpt_oss_cross_provider_failed_rows_1200
基本信息
- 数据集地址:https://huggingface.co/datasets/dinushiTJ/nz_research_commons_gpt_oss_cross_provider_failed_rows_1200
- 数据集大小:1,508,742 字节
- 下载大小:541,661 字节
- 数据切分:仅包含训练集(train),共 78 个样本
数据特征(共 25 个字段)
文本与元数据
title(string):标题authors(string):作者subjects(string):主题abstract(string):摘要text(string):文本内容record_id_hash(string):记录 ID 哈希值label_text(string):标签文本
提示与对话
prompt(string):提示conversations(string):对话内容
分类与置信度
classification_label(int64):分类标签classification_reason(string):分类原因classification_confidence(float64):分类置信度
嵌入信息
embedding_text(string):嵌入文本embedding_token_count(int64):嵌入 token 数量
GPT OSS 相关字段
gpt_oss_model(string):GPT OSS 模型名称gpt_oss_reasoning_effort(string):GPT OSS 推理努力程度gpt_oss_label(null):GPT OSS 标签(空值)gpt_oss_confidence(null):GPT OSS 置信度(空值)gpt_oss_reason(null):GPT OSS 原因(空值)gpt_oss_raw_output(string):GPT OSS 原始输出gpt_oss_error(string):GPT OSS 错误信息
其他
token_count(int64):token 数量year(string):年份row_index(int64):行索引seconds(float64):耗时(秒)
数据配置
- 配置名称:default
- 数据文件路径:
data/train-*(所有匹配该模式的文件)




