遇见数据集

A modelled dataset of compute demand and service requirements for Chinese provinces

收藏
Zenodo2026-06-08 更新2026-06-05 收录
官方服务:

资源简介:

This record contains a modelled, scenario-based dataset of fine-grained computing-power (compute) demand for the 31 provincial-level administrative regions of mainland China. The dataset is produced by a fully reproducible three-stage pipeline that deliberately decouples three independent sources of evidence: 1. Workload morphology. Temporal shape, online/batch workload composition and task-size distributions are extracted from the Huawei sir-lab 2023 private serverless trace (200 functions, per-minute resolution, 141 of 235 observed days). Functions are classified into online (latency-sensitive) and batch (throughput-oriented) classes; demand is reconstructed under two estimates — a conservative "basic" estimate from observed counts and an upper-bound "aggr" estimate from a right-censored Tobit/EM model — and the partial trace is extended to a full calendar year with a calendar-aware GAM-Poisson model. 2. National capacity anchoring. The reference curve is rescaled to the national installed computing capacity reported in the 2025 Advanced Computing and Computing Power Development Blue Paper (962 EFlops), under four demand scenarios at 60%, 100%, 150% and 200% of capacity (approximately 577, 962, 1443 and 1924 EFlops). A single scaling coefficient preserves timing, online/batch morphology and task-size distributions. 3. Provincial adjustment. The national baseline is redistributed across the 31 provinces using a size–shape decomposition driven by a multi-indicator socioeconomic panel (digital access, software/IT services, urban and service-sector scale, higher education and R&D, patents, public S&T expenditure). Provincial scale is set by entropy-weighted indicators and the provincial 24-hour shape by a composition of five activity patterns. Contents- Provincial compute-demand time series: online workloads at minute resolution (525,600 records per province) and batch workloads at hourly resolution (8,760 records per province), for both estimates (basic, aggr) and all four scenarios (s060, s100, s150, s200).- Service-requirement distributions per province, task type, day type and hour (72 buckets), in normalised core-seconds.- Provenance, scenario and provincial-adjustment metadata, and code to reproduce the full pipeline. A manifest with row counts and SHA-256 checksums is included for integrity checking. Technical validationNational aggregation reproduces the target scenario totals; province-level annual demand correlates with independent provincial peak electricity load (Spearman rho ~ 0.77; log-scale Pearson ~ 0.82); and the online/batch diurnal patterns reproduce the ~8-hour peak separation reported in cloud-workload studies. Intended use: research on data-centre and computing-infrastructure planning, regional energy/electricity-load coupling, and workload-aware scheduling.

提供机构:
Zenodo
创建时间:
2026-06-04
二维码
社区交流群
二维码
科研交流群
商业服务