MA-tokenweights/wikitext-103-raw-v2-tf-idf-wordlevel-sublinear
收藏资源简介:
该数据集是一个用于自然语言处理任务的文本数据集,包含文章级别的文本信息及其相关特征。数据集由268,336个训练样本、540个验证样本和618个测试样本组成,总大小约2.98 GB。特征包括文章ID、来源分割信息、文章文本内容、单词列表、TF-IDF分数、标记ID、标记权重、标记与单词的索引关系、标记是否为第一个或最后一个子标记等,适用于文本分析、关键词提取、词权重计算或模型训练等任务。
This dataset is a text dataset for natural language processing tasks, containing article-level text information and related features. It consists of 268,336 training samples, 540 validation samples, and 618 test samples, with a total size of approximately 2.98 GB. Features include article ID, source split information, article text content, word lists, TF-IDF scores, token IDs, token weights, token-to-word indices, and indicators for whether a token is the first or last subtoken, making it suitable for tasks such as text analysis, keyword extraction, token weight calculation, or model training.



