遇见数据集

NewsMinerCollection

收藏
Mendeley Data2019-10-24 更新2026-04-09 收录
官方服务:

资源简介:

The NewMinerCollection was built by collecting news, in the English language, from The Guardian, CNN, BBC, Fox News, NyPost, China Daily and CNBC websites, from 1990 until 2016 using a web crawler. This dataset contains 7000 news items equally distributed among the seven categories: id_category category 1 Arts, Culture & Entertainment 4 Economy, Business & Finance 6 Environmental Issues 10 Lifestyle & Leisure 11 Politics 15 Sport 13 Science & Technology

NewMinerCollection数据集通过网络爬虫(web crawler)于1990年至2016年间,从《卫报》(The Guardian)、美国有线电视新闻网(CNN)、英国广播公司(BBC)、福克斯新闻(Fox News)、《纽约邮报》(NyPost)、《中国日报》(China Daily)以及美国消费者新闻与商业频道(CNBC)的官方网站采集英文新闻构建而成。该数据集共包含7000条新闻,均匀分布于7个类别中,具体分类及对应编号如下:编号1为艺术、文化与娱乐,编号4为经济、商业与金融,编号6为环境议题,编号10为生活方式与休闲,编号11为政治,编号15为体育,编号13为科学与技术。

创建时间:
2019-10-24
二维码
社区交流群
二维码
科研交流群
商业服务