Index / koronavírus
收藏资源简介:
This object has been created as a part of the web harvesting project of the Eötvös Loránd University Department of Digital Humanities ELTE DH. Learn more about the workflow HERE about the software used HERE.The aim of the project is to make online news articles and their metadata suitable for research purposes. The archiving workflow is designed to prevent modification or manipulation of the downloaded content. The current version of the curated content with normalized formatting in standard TEI XML format with Schema.org encoded metadata is available HERE. The detailed description of the raw content is the following: The portal's archived content (from 2013-03-28 to 2021-01-30) in WARC format available HERE (crawled: 2021-01-30 13:01:33.164201 - 2021-01-31 19:47:27.191660). The portal's archived content (from 2021-01-31 to 2021-05-24) in WARC format available HERE (crawled: 2021-05-25 08:47:51.634068 - 2021-05-25 17:04:43.254877).
本数据集为厄特沃什·罗兰大学数字人文系(Eötvös Loránd University Department of Digital Humanities,简称ELTE DH)网络爬取项目的成果之一。有关工作流程与所用软件的详情,请点击此处查看。 本项目旨在使在线新闻文章及其元数据适配学术研究需求。本次存档工作流程的设计目标为防止下载的内容被篡改或非法操纵。 当前经整理的标准化格式内容采用标准TEI XML(Text Encoding Initiative XML)格式,并附带Schema.org编码的元数据,其最新版本可点击此处获取。 原始内容的详细说明如下: 该门户网站的存档内容(时间范围:2013年3月28日至2021年1月30日)以Web归档格式(WARC)存储,可点击此处获取(爬取时间:2021年1月30日13:01:33.164201 至 2021年1月31日19:47:27.191660)。 该门户网站的存档内容(时间范围:2021年1月31日至2021年5月24日)以WARC格式存储,可点击此处获取(爬取时间:2021年5月25日08:47:51.634068 至 2021年5月25日17:04:43.254877)。



