Economic Relevant News from The Guardian
收藏资源简介:
The news: The present dataset consists of 1789 news articles from the British daily newspaper The Guardian extracted using the content endpoint of The Guardian Open Platform. The news articles were, at the time, all the news corresponding to the sections: business, politics, society and world news for the entire month of January of 2013 (for a total of 1689 news) and an extra set of news articles randomly selected from the period Febrary of 2013 to December of 2015 (100 news articles). The first set of 1689 news articles was used for training and the second set of 100 news articles was used for testing in two publications: * Maisonnave, M., Delbianco, F., Tohmé, F.A. and Maguitman, A.G., 2018, November. A Supervised Term-Weighting Method and its Application to Variable Extraction from Digital Media. In XIX Simposio Argentino de Inteligencia Artificial (ASAI)-JAIIO 47 (CABA, 2018). * Maisonnave, M., Delbianco, F., Tohmé, F.A. and Maguitman, A.G., 2019. A Flexible Supervised Term-Weighting Technique and its Application to Variable Extraction and Information Retrieval. Inteligencia Artificial, 22(63), pp.61-80. The labels: The entire dataset was manually classified into two possible categories: economically relevant and irrelevant. The labelling process was carried out by two experts in Economy working in collaboration. For each news article, the full text of the article was analyzed to determine the category. The format: There are two different versions for this dataset: the reduced and the full versions. The former consists of a CSV and a readme file. The CSV file has five columns: "Instance No.", "Title", "Web Publication Date", "web URL" and "Economically Relevant". This version is reduced in columns as it does not include the full article texts; however, it does include all the 1789 instances. Requesting the full dataset: To gain access to the full version of the dataset (which includes the body of the news articles), please send an email to mariano.maisonnave@cs.uns.edu.ar with a copy to openplatform@theguardian.com requesting authorization and making it clear that the data set will not be used for commercial purposes.
新闻数据集: 本数据集包含1789篇来自英国日报《卫报》(The Guardian)的新闻文章,通过《卫报》开放平台(The Guardian Open Platform)的内容接口提取所得。所收录的新闻分为两部分:其一为2013年1月全月的商业、政治、社会与国际新闻板块的全部内容,共计1689篇;其二为2013年2月至2015年12月期间随机选取的额外100篇新闻。其中1689篇新闻用于模型训练,剩余100篇用于模型测试,相关研究成果已发表于以下两篇学术出版物: * Maisonnave, M.、Delbianco, F.、Tohmé, F.A. 与 Maguitman, A.G.,2018年11月。《一种监督式词权重方法及其在数字媒体变量抽取中的应用》,载于第47届阿根廷人工智能研讨会(XIX Simposio Argentino de Inteligencia Artificial, ASAI)暨第47届阿根廷信息学联合大会(JAIIO 47,布宜诺斯艾利斯,2018)。 * Maisonnave, M.、Delbianco, F.、Tohmé, F.A. 与 Maguitman, A.G.,2019。《一种灵活的监督式词权重技术及其在变量抽取与信息检索中的应用》,载于《人工智能》(Inteligencia Artificial)第22卷第63期,第61-80页。 标签体系: 本数据集经人工标注为两类:经济相关与非经济相关。标注工作由两位经济学领域专家协作完成,针对每篇新闻的全文进行分析以确定其类别。 数据集格式: 本数据集包含精简版与完整版两种形式。精简版包含一个CSV文件与一份说明文档(readme)。该CSV文件共包含五列:「实例编号(Instance No.)」、「标题(Title)」、「网络发布日期(Web Publication Date)」、「网络链接(Web URL)」与「经济相关性(Economically Relevant)」。该精简版未收录新闻全文,仅保留上述核心列字段,但涵盖全部1789条数据实例。 完整版数据集获取: 如需获取包含新闻正文的完整版数据集,请发送邮件至mariano.maisonnave@cs.uns.edu.ar,并抄送openplatform@theguardian.com,申请授权并明确说明本数据集将不用于商业用途。




