遇见数据集

Rijksmuseum Web Presence in the Internet Archive: Research Datasets 1999–2025

收藏
Zenodo2026-07-06 更新2026-08-02 收录
官方服务:

资源简介:

This deposit contains datasets and analysis notebooks on the web‑archived presence of the Rijksmuseum (rijksmuseum.nl) in the Internet Archive’s Wayback Machine. The material was developed in a broader research project on openness, web preservation, and museum web history, and is partly used in the article “How Open is the Rijksmuseum? Web Archives, Documentation, and the Historical Record of Online Collections” (Journal of Digital History) by Nadezhda Povroznik. The deposit includes: Raw Internet Archive exports for rijksmuseum.nl (1999–June 2025), as returned by the Wayback Machine API. A cleaned URL‑level dataset (status 200 only, digest‑based deduplication, removal of query parameters and fragments, yearly unique URLs). Derived summary tables, e.g. yearly URL counts, format distributions, and key URL segment and language statistics. A set of Jupyter notebooks documenting the main processing and analysis steps from raw data to the cleaned and derived datasets. The notebooks in this deposit form a frozen, citable snapshot that matches the data. A more up‑to‑date and extended version of the code and analyses is maintained in the project’s GitHub repository. These resources can be used to reproduce and extend the analyses, to study how a major museum is represented in web archives over time, and to explore methodological questions around URL‑based analysis and source criticism for web‑archived collections. Users should note that the data reflect only what the Internet Archive preserved - coverage is uneven, and dynamically generated content is only partially captured.

提供机构:
Zenodo
创建时间:
2026-07-06
二维码
社区交流群
二维码
科研交流群
商业服务