遇见数据集

Collection of web pages. For events. Coronavirus (COVID-19)

收藏
data.europa2024-07-03 收录
官方服务:

资源简介:

The collection of web pages is the main way to carry out the legal deposit of online publications. It is carried out with crawling robots that go through the previously selected URLs and save everything they have linked to the frequency, depth and size that is determined. The result of these web collections are web files. These are selective collections, they are made to complete mass collections, since they collect with greater depth and frequency a smaller sample of websites selected for their relevance to history, society and culture. They are held several times a year in collaboration with the conservation centres of the autonomous communities and other specialised institutions. Events gather on events of special relevance.

网页集合是开展在线出版物法定缴送的主要载体。该采集工作通过爬取机器人完成:机器人遍历预先选定的统一资源定位符(URL),并按照预设的采集频率、爬取深度与数据规模,保存其关联的全部内容。此类网页集合的产出产物为网页归档文件。 此类采集属于选择性采集,用于补充批量采集工作:针对与历史、社会及文化具有相关性的精选网站样本,以更高的采集深度与频率进行数据采集。该工作每年会同各自治社区的保护中心及其他专业机构协同开展多次。 此外,会针对具有特殊重要性的事件开展专项采集。

二维码
社区交流群
二维码
科研交流群
商业服务