遇见数据集

MA-tokenweights/all-the-news-2-tfidf-topic-stratified-v1-articles

收藏
Hugging Face2026-05-19 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个包含新闻文章或类似文本数据的集合,用于自然语言处理任务。数据集包括文章ID、源分割信息、文章文本内容、元数据(如标题、出版物、日期、URL、章节)以及NLP处理特征(如单词列表、TF-IDF分数、令牌ID和权重)。数据被分割为训练文章和验证文章两部分,适用于机器学习模型的训练和评估。

This dataset is a collection of news articles or similar text data designed for natural language processing tasks. It includes article IDs, source split information, article text content, metadata (such as title, publication, date, URL, section), and NLP processing features (e.g., word lists, TF-IDF scores, token IDs, and weights). The data is split into training articles and validation articles, making it suitable for training and evaluating machine learning models.

提供机构:
MA-tokenweights
二维码
社区交流群
二维码
科研交流群
商业服务