nz_research_commons_inference_results
收藏资源简介:
该数据集是一个用于文本分类任务的学术文献集合,专门用于识别文献是否具有毛利文化起源。数据集包含3318个训练样本,每个样本包含丰富的文献元数据和分类信息,核心特征包括:文献标题、作者、学科主题、摘要、全文内容、年份等基本信息;二分类标签(毛利起源/非毛利起源)及分类理由说明;文本长度统计(token计数);多种嵌入表示文本和模型输出结果,包括基础嵌入标签、ID、分数以及4B模型的响应文本。数据集适用于文化起源分类、文本分析、模型训练与评估等自然语言处理任务,特别关注毛利文化相关文献的识别与研究。
This dataset is an academic literature collection for text classification tasks, specifically designed to identify whether literature has Maori cultural origins. It contains 3318 training samples, each with rich metadata and classification information, including core features such as: basic information like title, author, subject, abstract, full text, and year; binary classification labels (Maori origin/non-Maori origin) and classification rationale; text length statistics (token count); and various embedded representations and model outputs, including basic embedding labels, IDs, scores, and responses from a 4B model. The dataset is suitable for natural language processing tasks such as cultural origin classification, text analysis, model training and evaluation, with a particular focus on the identification and study of literature related to Maori culture.
- 数据集名称:
nz_research_commons_inference_results - 数据集链接: https://huggingface.co/datasets/dinushiTJ/nz_research_commons_inference_results
- 数据集大小: 下载大小约16.21 MB,数据集总大小约43.79 MB
- 数据划分: 仅包含训练集(train),共3,318个样本
- 特征字段:
title: 标题(字符串)authors: 作者(字符串)subjects: 主题(字符串)abstract: 摘要(字符串)text: 文本内容(字符串)record_id_hash: 记录ID哈希(字符串)prompt: 提示词(字符串)token_count: 令牌数量(整数)classification_label: 分类标签,包含两个类别:0: maori_origin(毛利起源)1: non_maori_origin(非毛利起源)
classification_reason: 分类理由(字符串)year: 年份(字符串)classification_confidence: 分类置信度(浮点数)embedding_text: 嵌入文本(字符串)embedding_token_count: 嵌入令牌数量(长整数)BaseEmbGemZeroLabel: 基础嵌入GemZero标签(字符串)BaseEmbGemZeroID: 基础嵌入GemZero ID(长整数)BaseEmbGemZeroScore: 基础嵌入GemZero得分(双精度浮点数)Base4BResponse: 基础4B响应(字符串)




