Public subset of a photovoltaic predictive maintenance decision dataset with human-in-the-loop annotations
收藏资源简介:
This dataset provides a public, non-sensitive subset of a photovoltaic predictive maintenance decision dataset collected in the context of a utility-scale solar power plant operating in a Saharan environment. The dataset is designed to support reproducible research on artificial intelligence–based decision-making for predictive maintenance under uncertainty, operational constraints, and human-in-the-loop supervision. The published subset contains annotated maintenance decision scenarios derived from Internet-of-Things monitoring streams and expert assessments. It includes sensor-derived features, environmental indicators, fault descriptors, and decision-related variables, as well as human validation labels used to supervise and refine algorithmic recommendations. All raw signals, plant identifiers, proprietary operational parameters, and commercially sensitive information have been removed or anonymized prior to publication. This dataset is associated with a hybrid decision-support framework combining fuzzy logic reasoning, adaptive decision trees, and reinforcement learning with human-in-the-loop feedback. It has been used to evaluate decision latency, adequacy, and operational robustness in harsh desert conditions, as reported in the companion research article submitted to Engineering Applications of Artificial Intelligence. Contents A structured tabular dataset of annotated predictive maintenance scenarios (CSV and Excel formats) Derived and normalized features suitable for machine learning and decision modeling Human-in-the-loop validation labels reflecting expert maintenance decisions Documentation describing feature definitions and data structure Reproducibility While the complete industrial dataset remains confidential due to operational and contractual constraints, this public subset enables independent validation of the proposed decision-making methodology. The feature space, decision labels, and data schema are consistent with those used in the full-scale study, allowing replication of model training, benchmarking, and comparative analyses. Scripts for data preprocessing and experimental configuration are provided separately in the associated code repository. Intended use This dataset is intended for research and educational purposes, including: AI-based predictive maintenance Human-in-the-loop decision systems Explainable and adaptive artificial intelligence Streaming and edge-constrained decision-making in energy systems Limitations The dataset represents a curated subset and does not include raw time-series signals or full plant-scale operational data. Results obtained using this subset should therefore be interpreted as representative of the decision logic and learning mechanisms rather than absolute plant performance.



