MA-tokenweights/all-the-news-2-tfidf-invfreq-topic-stratified-v1-val-sequences
收藏资源简介:
该数据集是一个用于自然语言处理任务的结构化数据集,包含文本序列的标记化表示和相关标注信息。特征包括输入ID序列(input_ids)、填充掩码(padding_mask)、标签序列(labels)、真实标记(real_tokens)、文章ID(article_id)、文章唯一标识符(article_uid)、数据源分割(source_split)、主题ID(topic_id)、主题强度(topic_strength)、序列索引(sequence_index)、标记起始位置(token_start)、标记结束位置(token_end)以及标记权重(token_weights)。数据集仅包含验证分割(validation_sequences),共1940个示例,文件大小约55.8 MB,适用于模型验证或评估任务,可能用于文本分类、序列标注或主题分析等场景。
This is a structured dataset designed for natural language processing (NLP) tasks, encompassing tokenized representations of text sequences and associated annotation information. Its features include input ID sequences (input_ids), padding masks (padding_mask), label sequences (labels), real tokens (real_tokens), article IDs (article_id), unique article identifiers (article_uid), data source splits (source_split), topic IDs (topic_id), topic strengths (topic_strength), sequence indices (sequence_index), token start positions (token_start), token end positions (token_end), and token weights (token_weights). This dataset only contains the validation split (validation_sequences), with a total of 1940 examples, and has a file size of approximately 55.8 MB. It is suitable for model validation or evaluation tasks, and can be applied to scenarios such as text classification, sequence labeling, or topic analysis.




