arcanumsearch-wiki
收藏资源简介:
该数据集是一个结构化的文本集合,包含804个训练样本。每个样本由三个字段组成:LINK(链接或标识符)、TITLE(标题)和TEXT(文本内容),所有字段均为字符串类型。数据集总大小为1,807,922字节,下载大小为972,983字节。数据以默认配置提供,文件路径指向训练分割。基于字段名称,该数据集可能适用于信息检索、文本分析或内容分类等任务,但具体背景、来源和应用场景未在README中明确说明。
This dataset is a structured text collection containing 804 training samples. Each sample consists of three fields: LINK (link or identifier), TITLE (title), and TEXT (text content), all of which are string types. The total dataset size is 1,807,922 bytes, with a download size of 972,983 bytes. The data is provided in a default configuration, with the file path pointing to the training split. Based on the field names, the dataset may be suitable for tasks such as information retrieval, text analysis, or content classification, but specific background, source, and application scenarios are not explicitly stated in the README.
数据集概述
- 数据集名称: arcanumsearch-wiki
- 数据集来源: Hugging Face 数据集平台
- 数据集大小: 下载大小约 0.93 MB,数据集总大小约 1.72 MB
数据特征
该数据集包含三个字段:
- LINK: 字符串类型,表示链接地址
- TITLE: 字符串类型,表示标题
- TEXT: 字符串类型,表示文本内容
数据划分
数据集仅包含一个划分:
- 训练集 (train): 包含 804 个样本,占用约 1.72 MB 存储空间
配置信息
- 配置名称: default
- 训练数据文件: 存储路径为
data/train-*




