High-Frequency Water Quality Time-Series Dataset for WAWQI Forecasting
收藏资源简介:
This dataset contains high-frequency, multivariate time-series data collected from an active freshwater aquaculture lake at Telkom University, Bandung, Indonesia. The data was recorded using a multi-sensor monitoring node deployed over a four-day period at a nominal 15-second interval. The primary purpose of this dataset is to facilitate research in short-term temporal forecasting of the Weighted Arithmetic Water Quality Index (WAWQI) using machine learning algorithms. The dataset consists of 16,980 rows and includes seven physical-chemical parameters. Additionally, the pre-computed WAWQI score is provided for direct use as a forecasting target, along with two columns describing the timing and continuity of the record. Dataset Variables: 1. Temperature_C: Water temperature in degrees Celsius (°C)2. pH: Potential of Hydrogen (acid-base balance)3. DO_mgL: Dissolved Oxygen in milligrams per liter (mg/L)4. Turbidity_NTU: Turbidity in Nephelometric Turbidity Units (NTU)5. EC_uScm: Electrical Conductivity in microSiemens per centimeter (µS/cm)6. TDS_mgL: Total Dissolved Solids in milligrams per liter (mg/L)7. ORP_mV: Oxidation-Reduction Potential in millivolts (mV)8. WAWQI_Score: The computed index score9. Timestamp: Device clock in YYYY-MM-DD HH:MM:SS format, populated for 10,984 of the 16,980 rows; the README explains why the remainder is empty10. Block: Continuous recording block, 1 to 5. Lag features, rolling statistics and forecast targets must not cross a block boundary. Note for reproducibility: This file is the analysis dataset exactly as used in the corresponding research. No trimming is required. Note on version 1: This version supersedes version 1, which published a 23,502-row file under a different preprocessing specification and is not comparable row for row. Version 1 labelled Electrical Conductivity in mS/cm; the correct unit is µS/cm.



