遇见数据集

Replication Package for "Predictive Autoscaling in Kubernetes via Machine Learning Time Series Forecasting"

收藏
Zenodo2026-02-20 更新2026-05-26 收录
官方服务:

资源简介:

Note: Links do not work in Zenodo but only from the README.MD file when downloaded locally. Contents The repository contains the following: INSTALL: Thorough installation guide describing how to test the system and install necessary tools and components of the system. DATASETS: Contains all test datasets extracted from testing phase GroundTruth: Zero autoscaling enabled Baseline: Default Kubernetes Horizontal Pod Autoscaler enabled Intermediate Study: The proposed Autoscaler solution enabled agg_minute.csv: Training dataset used for training the models COMPONENTS: Contains the components and source code which constitute the solution Autoscaler: Backend and frontend for the solution Forecaster: Model training and prediction Workload: Workload generation for testing FIGURES: Contains all figures from the paper NB. Workload 0 defines the generated data which closely resembles the traning data, while workload 1 is the sinusoidally generated data. Paper figures: Ground Truth 10x10 workload 0 Ground Truth 10x10 workload 0 Baseline workload 0 Baseline workload 0 Autoscaler workload 0 Autoscaler workload 0 Architectural Diagram: Architectural diagram describing the proposed system Preprocessing pipeline: Preprossing pipeline flow Figures with startup-time infixed e.g intermediate-study-v2.2-startup-time-workload-0.pdf displays the extended tests using applications with longer startup times. MODELS: Contains all the baseline models and selected hyperparameters from tuning phase. The baseline models used are present in the p9 folder. RESULTS: Contains results from system tests extracted from datasets using scripts SCRIPTS: Contains scripts to generate figures plot_module.py: Reads csv, aggregates the input, formats figures and outputs to results in .svg, .png and .pdf format make_figures.py: Uses above module to iterate through files in --input_dir to create a plot for each input file. plot_training_data.py: Used to plot the training data figure from agg_minute.csv

注意:本仓库在Zenodo平台上的链接无法正常访问,仅在本地下载后可通过README.MD文件获取有效链接。 目录 本仓库包含以下内容: ### 安装指南(INSTALL):详细说明如何测试本系统并安装系统所需的工具与组件。 ### 数据集(DATASETS):包含测试阶段提取的全部测试数据集 - 真实值(GroundTruth):已启用零自动扩缩容(Zero autoscaling)的测试集 - 基线(Baseline):已启用默认Kubernetes水平Pod自动扩缩容器(Horizontal Pod Autoscaler, HPA)的测试集 - 中间研究(Intermediate Study):已启用本文提出的自动扩缩容解决方案的测试集 - agg_minute.csv:用于模型训练的训练数据集 ### 组件(COMPONENTS):包含构成本解决方案的全部组件与源代码 - 自动扩缩容器(Autoscaler):本解决方案的前后端代码 - 预测器(Forecaster):模型训练与预测模块 - 工作负载生成器(Workload):用于测试的工作负载生成模块 ### 图表(FIGURES):包含论文中的全部图表 备注:工作负载0(Workload 0)为生成的与训练数据特征高度相似的数据集,而工作负载1(Workload 1)为正弦曲线生成的数据集。 #### 论文图表: - 真实值(Ground Truth)10×10 工作负载0 - 真实值10×10 工作负载0 - 基线方案(Baseline)工作负载0 - 基线方案工作负载0 - 本文自动扩缩容方案工作负载0 - 本文自动扩缩容方案工作负载0 - 架构图:阐述本文所提出系统的架构示意图 - 预处理流水线:预处理流程示意图 文件名中嵌入启动时间信息的图表(如intermediate-study-v2.2-startup-time-workload-0.pdf)展示了针对启动耗时更长的应用程序开展的扩展测试。 ### 模型库(MODELS):包含全部基线模型与调优阶段筛选出的最优超参数。所用基线模型存储于p9文件夹中。 ### 结果集(RESULTS):包含通过脚本从数据集中提取的系统测试结果。 ### 脚本集(SCRIPTS):包含用于生成图表的脚本 - plot_module.py:读取CSV文件、聚合输入数据、格式化图表并输出.svg、.png及.pdf格式的结果文件 - make_figures.py:调用上述模块,遍历--input_dir指定目录下的文件,为每个输入文件生成对应图表 - plot_training_data.py:用于根据agg_minute.csv绘制训练数据图表

提供机构:
Zenodo
创建时间:
2026-02-20
二维码
社区交流群
二维码
科研交流群
商业服务