Comparative analysis of time series clustering methods : simulation data
收藏资源简介:
Simulated data: construction of the reference clusters This dataset contains the synthetic incidence series used to build the nine reference clusters against which the clustering methods were evaluated. Each series was constructed as two consecutive epidemic waves, each wave drawn from one of the three transmission scenarios (R₀ = 11.52, 2.88 or 1.06). The simulated trajectories were taken from the SEIR simulation output of Barro et al. [10.5281/zenodo.22554829], in which 8,900 trajectories were retained per scenario after a priori filtering. The nine possible ordered combinations of a first and a second wave define the nine reference clusters (simulated clusters SC1 to SC9), each combination being labelled by its (first-wave R₀ → second-wave R₀) pair: SC1: 11.52 → 11.52; SC2: 11.52 → 2.88; SC3: 11.52 → 1.06 SC4: 2.88 → 2.88; SC5: 2.88 → 11.52; SC6: 2.88 → 1.06 SC7: 1.06 → 1.06; SC8: 1.06 → 11.52; SC9: 1.06 → 2.88 For each of the three scenarios, 1,000 series were randomly drawn from the pool of 8,900 valid trajectories. This sample size was chosen for two reasons: to limit the substantial computational burden of the functional analysis, and to approximate real surveillance conditions, in which the number of monitored spatial units is limited. The nine reference clusters, each comprising 1,000 series, served as the ground-truth partition for evaluating the performance of the clustering methods.



