Document Type : Research Article
Authors
Department of Irrigation and Reclamation Engineering, College of Agriculture and Natural Resources, University of Tehran, Karaj, Iran
Abstract
Introduction
Precipitation is a fundamental component of the hydrological cycle and plays a vital role in water resource management, agriculture, ecosystem sustainability, hydropower generation, disaster mitigation, and urban planning. Accurate rainfall prediction is essential for effective water allocation, flood and drought risk management, and the development of early warning systems, thereby reducing potential damages to human communities and infrastructure. In agriculture, reliable precipitation forecasts support optimal planting schedules and resource management, leading to increased productivity and reduced weather-related losses. Furthermore, rainfall prediction is crucial for reservoir operation and sustainable long-term water resource planning.
Due to the limited spatial coverage of ground-based meteorological stations, satellite-based precipitation products have gained increasing attention as reliable alternatives for regional-scale rainfall analysis. Among these, the Climate Hazards Group InfraRed Precipitation with Stations (CHIRPS) dataset is widely recognized for its extensive spatial coverage, relatively high accuracy in arid and mountainous regions, and free accessibility. CHIRPS integrates satellite observations with ground station data and provides precipitation estimates at approximately 5 km spatial resolution across multiple temporal scales since 1981. In this study, monthly mean CHIRPS precipitation data for the period 2000–2024 were extracted and analyzed for rainfall prediction purposes.
The inherent nonlinearity and complex behavior of precipitation processes make traditional statistical approaches insufficient for accurate forecasting. Consequently, machine learning techniques have emerged as powerful tools for modeling complex climatic phenomena. Although numerous previous studies have applied machine learning algorithms, such as Artificial Neural Networks (ANN), Random Forest (RF), Support Vector Machines (SVM), and Convolutional Neural Networks (CNN), to rainfall prediction, most have focused on single-model evaluations or simple model comparisons. However, combining multiple models has been shown to reduce prediction uncertainty and improve forecast robustness, an approach that has received comparatively limited attention.
The primary objective of this research is to evaluate the performance of satellite-based CHIRPS precipitation data and to enhance monthly rainfall prediction using machine learning models implemented in the WEKA environment. To achieve this goal, fifteen individual machine learning models were first assessed independently. Subsequently, pairwise combinations of these models (105 combinations in total) were developed using simple averaging and brute-force weighting strategies to improve predictive accuracy. The novelty of this study lies not merely in the application of multiple machine learning algorithms but in proposing a systematic, transparent, and reproducible framework for satellite-based rainfall prediction in data-scarce basins. This framework is region-independent and can be readily applied to other climatic and hydrological settings.
Materials and Methods
This study investigates monthly precipitation prediction in the Salt Lake Basin using satellite-based CHIRPS data for the period January 2000 to December 2025, extracted via the Google Earth Engine platform. CHIRPS was selected due to its adequate spatial resolution (0.05°), integration of satellite and ground-based observations, long-term temporal coverage since 1981, and suitability for arid and semi-arid regions with sparse rain-gauge networks. Data quality was assessed using the Interquartile Range (IQR) method, and no outliers were detected. As the dataset exhibited consistent scaling across all time steps, no normalization was applied. The data were divided into training (January 1999–December 2021) and testing (January 2022–December 2024) subsets. All modeling procedures were implemented in the WEKA environment using time-series modules. To validate the CHIRPS dataset, satellite-derived precipitation estimates were compared with observations from four ground stations over ten years (2007–2017), and basin-scale performance was evaluated using averaged statistical metrics. Initially, sixteen machine learning models were examined, and one model (Random Tree) was excluded due to poor performance. The final set of fifteen models included Gaussian Processes, Random Forest, MLP Regressor, RBF Regressor, Linear Regression, SMOreg, IBk, LWL, Additive Regression, Bagging, Random Committee, Random SubSpace, Decision Table, M5Rules, and M5P. Model hyperparameters were optimized using the Grid Search tool in WEKA. To improve prediction accuracy and robustness, pairwise ensemble models were generated using simple averaging, resulting in 105 combined configurations. Additionally, a brute-force weighting approach was applied by assigning weights between 0 and 1 with a step of 0.1 to each model, subject to a unity-sum constraint, in order to identify the optimal ensemble structure.
Model performance was evaluated using the correlation coefficient, RMSE, MAE, MSE, bias, and Nash–Sutcliffe efficiency (NSE), and the most accurate configuration was selected based on testing results.
Results and Discussion
The accuracy of the CHIRPS satellite precipitation product was first evaluated using observed rainfall data from four rain-gauge stations distributed across the Salt Lake Basin. The validation results indicated a satisfactory agreement between satellite-derived and observed precipitation, with an overall coefficient of determination (R²) of approximately 0.69 and an NSE of 0.70, confirming the suitability of CHIRPS data for monthly rainfall analysis in the study area. Subsequently, the performance of fifteen individual machine learning models was assessed for monthly precipitation prediction using the testing period (January 2022–December 2024). Model evaluation based on correlation coefficient, RMSE, MAE, MSE, bias, and Nash–Sutcliffe efficiency (NSE) revealed notable differences among algorithms. Rule-based and tree-based models, particularly M5Rules and Additive Regression, exhibited superior performance, characterized by higher NSE values, lower error magnitudes, and relatively small bias. In contrast, instance-based models such as IBk and function-based models like RBF Regressor showed weaker performance, with larger error dispersion and lower correlation with observed data. The combined analysis of bias, residual standard deviation, and median absolute error provided a more comprehensive understanding of model behavior than single error metrics alone. Visual assessments using Taylor diagrams and scatter plots further confirmed the robustness of M5Rules, followed by Additive Regression and Random Forest, in capturing both the variability and temporal patterns of observed precipitation.
To enhance prediction accuracy and stability, pairwise combinations of the individual models were developed using simple averaging, resulting in 105 ensemble configurations. The results demonstrated that most ensemble models outperformed their corresponding single-model counterparts, as evidenced by higher NSE values and reduced error metrics. The combination of Additive Regression–M5Rules achieved the best overall performance, yielding the lowest MAE and MSE and the highest NSE among all tested configurations. Histogram analysis of NSE values showed that a large proportion of the ensemble models achieved NSE values above 0.75, indicating the general effectiveness of the ensemble approach rather than improvement limited to specific combinations. In addition, a brute-force weighting strategy was applied to optimize model contributions within the ensembles. Compared to simple averaging, the Brute Force approach consistently improved NSE values by reducing the influence of weaker models and assigning higher weights to more accurate ones. This improvement was particularly evident in combinations involving low-performing models such as IBk and RBF Regressor. Nevertheless, simple averaging also produced competitive and stable results without increasing computational complexity.
Overall, the results indicate that monthly rainfall prediction accuracy depends not only on the choice of individual algorithms but also on the interaction and complementarity of model errors. The proposed ensemble framework effectively reduced bias and variance while avoiding overfitting through strict temporal data separation and independent testing. These findings demonstrate that systematic pairwise model combination, even with simple averaging, can substantially enhance the accuracy and robustness of precipitation forecasts in data-scarce and climatically complex regions.
Conclusion
The Salt Lake Basin is characterized by sparse rain-gauge coverage and high climatic variability, making accurate monthly precipitation estimation essential for water resources management. This study evaluated the CHIRPS satellite precipitation dataset and assessed the performance of individual and combined machine learning models for monthly rainfall prediction. Validation against four ground stations confirmed that CHIRPS reliably represents the temporal pattern of monthly precipitation in the basin, with acceptable correlation and Nash–Sutcliffe efficiency values, indicating its suitability for hydrological applications in arid and semi-arid regions. Among the fifteen evaluated machine learning models, M5Rules, Additive Regression, and Random Forest exhibited superior performance as single models. Ensemble modeling further improved prediction accuracy by reducing systematic errors and enhancing robustness. The Additive Regression–M5Rules combination achieved the best performance (NSE ≈ 0.88), outperforming all individual models. These results highlight the effectiveness of integrating CHIRPS data with ensemble machine learning approaches for improving monthly precipitation prediction in data-scarce basins. The proposed framework provides a reliable and transferable tool for supporting sustainable water resources management in arid and semi-arid regions.
Keywords
Subjects