遇见数据集

News Mining Dataset for Sentiment and Topic Analysis: 300K Articles Extracted across 20 News Sources

收藏
Zenodo2025-04-16 更新2026-05-26 收录
官方服务:

资源简介:

This dataset was created using Arquivo.pt, the Portuguese web archive, as the primary source for extracting and analyzing news-related links. Over 3 million archived URLs were collected and processed, resulting in a curated collection of approximately 300,000 high-quality news articles from 20 different news sources. Each article has been processed to extract key information such as: Publication date News source Mentioned topics Sentiment analysis The dataset was built as part of a web-based application for relationship detection and exploratory analysis of news content. It can support research in areas such as natural language processing (NLP), computational journalism, network analysis, topic modeling, and sentiment tracking. All articles are in Portuguese, and the dataset is structured for easy use with tools like Python (e.g., Pandas, Spark) and machine learning workflows. Dataset Structure The dataset consists of two main folders: news/Contains all ~3 million processed URLs, organized by folders based on processing status: success/ — articles successfully extracted duplicated/ — duplicate content detected not_news/ — filtered out as non-news error/ — extraction or parsing failuresEach subfolder contains JSON files, partitioned as outputted by Spark. These represent the raw extracted content. news_processed/Contains 8 Parquet files, which are partitions of a cleaned and enriched dataset with approximately 300,000 high-quality news articles. These include structured fields ready for analysis.

提供机构:
Zenodo
创建时间:
2025-04-16
二维码
社区交流群
二维码
科研交流群
商业服务