遇见数据集

Historical UCL COVID-19 statistics

收藏
Zenodo2026-07-24 更新2026-08-02 收录
官方服务:

资源简介:

Between October 2020 and May 2022, University College London published statistics of COVID-19 cases among its staff and students on a web page that was overwritten with each update, keeping no history. This dataset is the result of a cron job that fetched that page hourly and extracted the numbers, so the series can be read as a whole rather than a day at a time. The original page has since been decommissioned and its URL redirected. The published data runs 2020-10-09 to 2022-05-11 (UCL's stated "final update Thursday 12 May 2022"); the 6,140 hourly page snapshots run to 2022-07-29. The deposit also includes 168 UCL COVID-19 update newsletters (9 March 2020 – 4 May 2022), the scraper and an archival re-implementation of it, a SHA-256 manifest of every fetch, and a self-contained interactive visualisation. To interpret the data, read README.md; to check its provenance, read PROVENANCE.md. Both are in the archive. They document the smoothing applied, what "on campus" meant, two gaps UCL never corrected, and other features that would otherwise look like errors. Licensing is mixed. The code and the extracted data (the CSV and JSON files) are published under the Apache License 2.0. The original UCL web pages in data/original/ and the UCL update newsletters in data/updates/ are copyright University College London: they are not covered by the Apache licence and are not relicensed, but reproduced as the evidence behind the extracted figures. Zenodo records a single licence field, which reflects only the Apache-licensed portion.

提供机构:
Zenodo
创建时间:
2026-07-24
二维码
社区交流群
二维码
科研交流群
商业服务