遇见数据集

PipeTwin

收藏
Zenodo2026-07-17 更新2026-08-02 收录
官方服务:

资源简介:

PipeTwin is a research framework that generates synthetic machine learning pipeline execution logs, builds a digital twin to model pipeline states, and uses explainable meta-learning to predict bottlenecks, failures, and resource usage. It automatically recommends optimal configurations for CPU, memory, batch size, and hyperparameters. 🚀 Key Features Synthetic Benchmark Generator – Generate up to 500 million observations with 120+ features, exported as highly compressed Parquet files. Digital Twin Engine – Continuously models pipeline state, resource utilization, and execution behavior. Explainable Meta-Learning – Uses ensemble learning (Random Forest, XGBoost, LightGBM, CatBoost, Extra Trees, etc.) to predict runtime, failures, bottlenecks, CPU, memory, and storage utilization. Explainable AI – SHAP, LIME, and Counterfactual explanations provide transparent recommendations. Autonomous Optimization – Automatically recommends CPU, memory, batch size, parallelism, caching strategy, and hyperparameter adjustments. Publication-Ready Visualizations – Generates high-quality PNG, PDF, and SVG figures at 300 DPI.

提供机构:
Zenodo
创建时间:
2026-07-16
二维码
社区交流群
二维码
科研交流群
商业服务