دوماهنامه

شبیه‌سازی بارش ماهانه با مدل‌های ترکیبی یادگیری ماشین: حوضه دریاچه نمک

نوع مقاله : مقالات پژوهشی

نویسندگان

گروه مهندسی آبیاری و آبادانی، دانشکدگان کشاورزی و منابع طبیعی دانشگاه تهران، کرج، ایران

چکیده
پیش‌بینی دقیق بارش ماهانه به‌ویژه در مناطق خشک و نیمه‌خشک، برای مدیریت منابع آب و برنامه‌ریزی‌های هیدرولوژیک اهمیت بالایی دارد. با توجه به کمبود و پراکندگی ایستگاه‌های زمینی در حوضه دریاچه نمک، استفاده از داده‌های ماهواره‌ای با دقت مکانی مناسب، ضرورت می‌یابد. نوآوری اصلی این پژوهش، ارزیابی کارایی داده‌های بارش CHIRPS در مدل‌سازی بارش ماهانه با بهره‌گیری از طیفی متنوع از مدل‌های یادگیری ماشین و همچنین بررسی اثربخشی مدل‌های ترکیبی در افزایش دقت پیش‌بینی است. در این مطالعه، داده‌های بارش CHIRPS با تفکیک افقی °05/0 برای دوره 2000 تا 2024 استخراج و پس از اعتبارسنجی اولیه مورد استفاده قرار گرفت. در این پژوهش برای ارزیابی داده‌های بارش CHIRPS، طیفی از مدل‌های یادگیری ماشین شامل مدل‌های درختی، مدل‌های رگرسیونی، مدل‌های کرنلی، مدل‌های مبتنی بر فاصله و مدل‌های ترکیبی به‌منظور مدل‌سازی رفتار غیرخطی بارش ماهانه مورد استفاده قرار گرفتند. این تنوع رویکردها، امکان ارزیابی تطبیقی ساختارهای مختلف یادگیری و بررسی تأثیر معماری مدل بر دقت پیش‌بینی را فراهم می‌سازد. مدل M5Rules با بدست آوردن 80/0 =  NSEبهترین عملکرد را در مدل‌های منفرد داشت. همچنین برای بهبود دقت، بر پایه میانگین‌گیری ساده، ۱۰۵ مدل ترکیبی به‌صورت دو‌به‌دو تولید و ارزیابی شد و سپس برای پیداکردن بهترین وزن از روش bruteforce (جستجوی فراگیر برای وزن‌دهی بهینه) استفاده شد. نتایج نشان داد که؛ مدل‌های ترکیبی به‌طور کلی عملکرد بهتری نسبت به مدل‌های منفرد دارند. ترکیب Additive Regression–M5Rules بهترین عملکرد را نسبت به سایر ترکیب‌ها با 155/0 =MAE ، 05/0 =MSE و 88/0=  NSEارئه کرد. این یافته‌ها نشان می‌دهد که ترکیب داده‌های CHIRPS با مدل‌های ترکیبی یادگیری ماشین می‌تواند رویکردی کارآمد و قابل‌ اتکا برای پیش‌بینی بارش ماهانه در مناطق با شبکه ایستگاهی محدود باشد.

کلیدواژه‌ها

موضوعات

عنوان مقاله English

Monthly Rainfall Simulation Using Hybrid Machine Learning Models: NAMAK Lake Basin

نویسندگان English

A. DalirGabrabad
M.A. Abdollahi
A. Ashrafzadeh
Department of Irrigation and Reclamation Engineering, College of Agriculture and Natural Resources, University of Tehran, Karaj, Iran
چکیده English

Introduction
Precipitation is a fundamental component of the hydrological cycle and plays a vital role in water resource management, agriculture, ecosystem sustainability, hydropower generation, disaster mitigation, and urban planning. Accurate rainfall prediction is essential for effective water allocation, flood and drought risk management, and the development of early warning systems, thereby reducing potential damages to human communities and infrastructure. In agriculture, reliable precipitation forecasts support optimal planting schedules and resource management, leading to increased productivity and reduced weather-related losses. Furthermore, rainfall prediction is crucial for reservoir operation and sustainable long-term water resource planning.
Due to the limited spatial coverage of ground-based meteorological stations, satellite-based precipitation products have gained increasing attention as reliable alternatives for regional-scale rainfall analysis. Among these, the Climate Hazards Group InfraRed Precipitation with Stations (CHIRPS) dataset is widely recognized for its extensive spatial coverage, relatively high accuracy in arid and mountainous regions, and free accessibility. CHIRPS integrates satellite observations with ground station data and provides precipitation estimates at approximately 5 km spatial resolution across multiple temporal scales since 1981. In this study, monthly mean CHIRPS precipitation data for the period 2000–2024 were extracted and analyzed for rainfall prediction purposes.
The inherent nonlinearity and complex behavior of precipitation processes make traditional statistical approaches insufficient for accurate forecasting. Consequently, machine learning techniques have emerged as powerful tools for modeling complex climatic phenomena. Although numerous previous studies have applied machine learning algorithms, such as Artificial Neural Networks (ANN), Random Forest (RF), Support Vector Machines (SVM), and Convolutional Neural Networks (CNN), to rainfall prediction, most have focused on single-model evaluations or simple model comparisons. However, combining multiple models has been shown to reduce prediction uncertainty and improve forecast robustness, an approach that has received comparatively limited attention.
The primary objective of this research is to evaluate the performance of satellite-based CHIRPS precipitation data and to enhance monthly rainfall prediction using machine learning models implemented in the WEKA environment. To achieve this goal, fifteen individual machine learning models were first assessed independently. Subsequently, pairwise combinations of these models (105 combinations in total) were developed using simple averaging and brute-force weighting strategies to improve predictive accuracy. The novelty of this study lies not merely in the application of multiple machine learning algorithms but in proposing a systematic, transparent, and reproducible framework for satellite-based rainfall prediction in data-scarce basins. This framework is region-independent and can be readily applied to other climatic and hydrological settings.
 
Materials and Methods
This study investigates monthly precipitation prediction in the Salt Lake Basin using satellite-based CHIRPS data for the period January 2000 to December 2025, extracted via the Google Earth Engine platform. CHIRPS was selected due to its adequate spatial resolution (0.05°), integration of satellite and ground-based observations, long-term temporal coverage since 1981, and suitability for arid and semi-arid regions with sparse rain-gauge networks. Data quality was assessed using the Interquartile Range (IQR) method, and no outliers were detected. As the dataset exhibited consistent scaling across all time steps, no normalization was applied. The data were divided into training (January 1999–December 2021) and testing (January 2022–December 2024) subsets. All modeling procedures were implemented in the WEKA environment using time-series modules. To validate the CHIRPS dataset, satellite-derived precipitation estimates were compared with observations from four ground stations over ten years (2007–2017), and basin-scale performance was evaluated using averaged statistical metrics. Initially, sixteen machine learning models were examined, and one model (Random Tree) was excluded due to poor performance. The final set of fifteen models included Gaussian Processes, Random Forest, MLP Regressor, RBF Regressor, Linear Regression, SMOreg, IBk, LWL, Additive Regression, Bagging, Random Committee, Random SubSpace, Decision Table, M5Rules, and M5P. Model hyperparameters were optimized using the Grid Search tool in WEKA. To improve prediction accuracy and robustness, pairwise ensemble models were generated using simple averaging, resulting in 105 combined configurations. Additionally, a brute-force weighting approach was applied by assigning weights between 0 and 1 with a step of 0.1 to each model, subject to a unity-sum constraint, in order to identify the optimal ensemble structure.
Model performance was evaluated using the correlation coefficient, RMSE, MAE, MSE, bias, and Nash–Sutcliffe efficiency (NSE), and the most accurate configuration was selected based on testing results.
 
Results and Discussion
The accuracy of the CHIRPS satellite precipitation product was first evaluated using observed rainfall data from four rain-gauge stations distributed across the Salt Lake Basin. The validation results indicated a satisfactory agreement between satellite-derived and observed precipitation, with an overall coefficient of determination (R²) of approximately 0.69 and an NSE of 0.70, confirming the suitability of CHIRPS data for monthly rainfall analysis in the study area. Subsequently, the performance of fifteen individual machine learning models was assessed for monthly precipitation prediction using the testing period (January 2022–December 2024). Model evaluation based on correlation coefficient, RMSE, MAE, MSE, bias, and Nash–Sutcliffe efficiency (NSE) revealed notable differences among algorithms. Rule-based and tree-based models, particularly M5Rules and Additive Regression, exhibited superior performance, characterized by higher NSE values, lower error magnitudes, and relatively small bias. In contrast, instance-based models such as IBk and function-based models like RBF Regressor showed weaker performance, with larger error dispersion and lower correlation with observed data. The combined analysis of bias, residual standard deviation, and median absolute error provided a more comprehensive understanding of model behavior than single error metrics alone. Visual assessments using Taylor diagrams and scatter plots further confirmed the robustness of M5Rules, followed by Additive Regression and Random Forest, in capturing both the variability and temporal patterns of observed precipitation.
To enhance prediction accuracy and stability, pairwise combinations of the individual models were developed using simple averaging, resulting in 105 ensemble configurations. The results demonstrated that most ensemble models outperformed their corresponding single-model counterparts, as evidenced by higher NSE values and reduced error metrics. The combination of Additive Regression–M5Rules achieved the best overall performance, yielding the lowest MAE and MSE and the highest NSE among all tested configurations. Histogram analysis of NSE values showed that a large proportion of the ensemble models achieved NSE values above 0.75, indicating the general effectiveness of the ensemble approach rather than improvement limited to specific combinations. In addition, a brute-force weighting strategy was applied to optimize model contributions within the ensembles. Compared to simple averaging, the Brute Force approach consistently improved NSE values by reducing the influence of weaker models and assigning higher weights to more accurate ones. This improvement was particularly evident in combinations involving low-performing models such as IBk and RBF Regressor. Nevertheless, simple averaging also produced competitive and stable results without increasing computational complexity.
Overall, the results indicate that monthly rainfall prediction accuracy depends not only on the choice of individual algorithms but also on the interaction and complementarity of model errors. The proposed ensemble framework effectively reduced bias and variance while avoiding overfitting through strict temporal data separation and independent testing. These findings demonstrate that systematic pairwise model combination, even with simple averaging, can substantially enhance the accuracy and robustness of precipitation forecasts in data-scarce and climatically complex regions.
Conclusion
The Salt Lake Basin is characterized by sparse rain-gauge coverage and high climatic variability, making accurate monthly precipitation estimation essential for water resources management. This study evaluated the CHIRPS satellite precipitation dataset and assessed the performance of individual and combined machine learning models for monthly rainfall prediction. Validation against four ground stations confirmed that CHIRPS reliably represents the temporal pattern of monthly precipitation in the basin, with acceptable correlation and Nash–Sutcliffe efficiency values, indicating its suitability for hydrological applications in arid and semi-arid regions. Among the fifteen evaluated machine learning models, M5Rules, Additive Regression, and Random Forest exhibited superior performance as single models. Ensemble modeling further improved prediction accuracy by reducing systematic errors and enhancing robustness. The Additive Regression–M5Rules combination achieved the best performance (NSE ≈ 0.88), outperforming all individual models. These results highlight the effectiveness of integrating CHIRPS data with ensemble machine learning approaches for improving monthly precipitation prediction in data-scarce basins. The proposed framework provides a reliable and transferable tool for supporting sustainable water resources management in arid and semi-arid regions.
 

کلیدواژه‌ها English

Climate data prediction
Precipitation data modeling
WEKA

Authors retain the copyright. This is an open access article distributed under Creative Commons Attribution 4.0 International License (CC BY 4.0).

  1. Abdollahi, M., Abedi, K.J., & Matinzadeh, M. (2024). The effect of the sub-basins area and methods of calculating concentration time on the simulation of urban runoff volume using SewerGEMS software (Case study: Shahrekord). https://doi.org/47176/jwss.28.3.49834
  2. Abinaya, P., & Janani, N. (2020). Rainfall forecasting using Weka data mining tool. International Research Journal of Engineering and Technology, 7(03).
  3. Abishek, B., Priyatharshini, R., Eswar, M.A., & Deepika, P. (2017). Prediction of effective rainfall and crop water needs using data mining techniques. 2017 IEEE Technological Innovations in ICT for Agriculture and Rural Development (TIAR). https://doi.org/10.1109/TIAR.2017.8273722
  4. Amrehn, M., Mualla, F., Angelopoulou, E., Steidl, S., & Maier, A. (2018). The random forest classifier in WEKA: Discussion and new developments for imbalanced data. arXiv preprint arXiv:1812.08102. https://doi.org/10.48550/arXiv.1812.08102
  5. Beck, H.E., Pan, M., Roy, T., Weedon, G.P., Pappenberger, F., Van Dijk, A.I., Huffman, G.J., Adler, R.F., & Wood, E.F. (2019). Daily evaluation of 26 precipitation datasets using Stage-IV gauge-radar data for the CONUS. Hydrology and Earth System Sciences, 23(1), 207–224. https://doi.org/10.5194/hess-23-207-2019, 2019
  6. Babaei Hessar, S., & Ghazavi, R. (2014). Comparison of TS and ANN models with the results of emission scenarios in rainfall prediction. Journal of Soil and Water. https://doi.org/10.22067/jsw.v0i0.10261
  7. Biau, G., & Scornet, E. (2016). A random forest guided tour. Test, 25(2), 197–227. https://doi.org/10.1007/s11749-016-0481-7
  8. Cervantes, J., Garcia-Lamont, F., Rodríguez-Mazahua, L., & Lopez, A. (2020). A comprehensive survey on support vector machine classification: Applications, challenges and trends. Neurocomputing, 408, 189–215. https://doi.org/ 10.1016/j.neucom.2019.10.118
  9. Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). https://doi.org/ 10.1145/2939672.2939785
  10. Dong, X., Yu, Z., Cao, W., Shi, Y., & Ma, Q. (2020). A survey on ensemble learning. Frontiers of Computer Science, 14(2), 241–258. https://doi.org/10.1007/s11704-019-8208-z
  11. Fürnkranz, J., Gamberger, D., & Lavrač, N. (2012). Foundations of Rule Learning. Springer. https://doi.org/10.1007/ 978-3-540-75197-7
  12. Fagan, M.E., & DeFries, R.S. (2024). Remote sensing and image processing. In S. M. Scheiner (Ed.), Encyclopedia of Biodiversity (3rd ed., pp. 432–445). Academic Press. https://doi.org/10.1016/B978-0-12-822562-2.00060-8
  13. Fallah-Ghalhari, Gh.A., Mousavi Baygi, M., & Habibi Nokhandan, M. (2009). Results compression of Mamdani fuzzy interface system and artificial neural networks in the seasonal rainfall prediction (Case study: Khorasan region).
  14. Funk, C., Peterson, P., Landsfeld, M., Pedreros, D., Verdin, J., Shukla, S., Husak, G., Rowland, J., Harrison, L., & Hoell, A. (2015). The climate hazards infrared precipitation with stations—a new environmental record for monitoring extremes. Scientific Data, 2(1), 1–21. https://doi.org/10.1038/sdata.2015.66
  15. Fikileni, S., & Wolski, P. (2022). Framework for implementation of the Pitman-WR2012 model in seasonal hydrological forecasting: A case study of Kraai River, South Africa. Water SA, 48(1), 62–74. https://doi.org/10.17159/WSA/2022.V48.I1.3891
  16. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press. https://doi.org/10.4258/hir. 2016.22.4.351
  17. Ghosh, S., Gourisaria, M.K., Sahoo, B., & Das, H. (2023). A pragmatic ensemble learning approach for rainfall prediction. Discover Internet of Things, 3(13). https://doi.org/10.1007/s43926-023-00044-3
  18. Haddadi, A., & Yazdi, M. (2025). Performance assessment of the extreme gradient boosting machine model in forecasting precipitation depth to improve precipitation estimation accuracy in data-scarce regions. Journal of Irrigation and Water Engineering. https://doi.org/10.22034/iwrr.2025.543507.2944
  19. Holmes, G., Hall, M., & Frank, E. (1999). Generating rule sets from model trees. Australian Joint Conference on Artificial Intelligence (pp. 1–12). Springer. https://doi.org/10.1007/3-540-46695-9_1
  20. James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An Introduction to Statistical Learning: with Applications in R (2nd ed.). Springer. https://doi.org/10.1007/978-1-4614-7138-7
  21. Kumar, P.S., & Swathi, M.N. (2023). Rainfall prediction using ada boost machine learning ensemble algorithm. Journal of Advanced Applied Scientific Research, 5(4), 67–81. https://doi.org/10.46947/joaasr542023682
  22. Kuncheva, L.I. (2014). Combining Pattern Classifiers: Methods and Algorithms (2nd ed.). John Wiley & Sons. https://doi.org/10.1109/TNN.2007.897478
  23. Mahjoobi, E., & Rafiei, F. (2025). Exploring the potential of geospatial data: An in-depth. In Exploring Remote Sensing-Methods and Applications (p. 113). https://doi.org/10.5772/intechopen.1006999
  24. Lin, S.-S., Zhu, K.-Y., Wang, C.-Y., Yang, C.-P., & Liu, M.-Y. (2025). Ensemble Learning-Based Soft Computing Approach for Future Precipitation Analysis. Atmosphere, 16(6), 669. https://doi.org/10.3390/atmos16060669
  25. Oswal, N. (2019). Predicting rainfall using machine learning techniques. arXiv preprint arXiv:1910.13827. https://doi.org/10.48550/arXiv.1910.13827
  26. Omary, A., Wedyan, A., Zghoul, A., Banihani, A., & Alsmadi, I. (2012). An interactive predictive system for weather forecasting. 2012 International Conference on Computer, Information and Telecommunication Systems (CITS). https://doi.org/10.1109/CITS.2012.6220375
  27. Patel, A., Keriwala, N., Soni, N., Goel, U., Bhoj, R., Adhyaru, Y., & Yadav, S. (2023). Rainfall prediction using machine learning techniques for Sabarmati River Basin, Gujarat, India. Journal of Engineering Science & Technology Review, 16(1). https://doi.org/10.25103/jestr.161.13
  28. Ramani, K., Reddy, M.S., Bhavani, K., Feeza, S., & Bavesh, V.S. (2024). Optimization of rainfall prediction using satellite data through machine learning and deep learning algorithms. 2024 IEEE International Conference on Information Technology, Electronics and Intelligent Communication Systems (ICITEICS). https://doi.org/10.1109/ ICITEICS61368.2024.10625624
  29. Rezaei, M., Nohtani, M., Moghaddamnia, A., Abkar, A., & Rezaei, M. (2014). Performance evaluation of Statistical Downscaling Model (SDSM) in forecasting precipitation in two arid and hyper arid regions. Journal of Soil and Water. https://doi.org/10.22067/jsw.v0i0.23119
  30. Rahman, A.-u., Abbas, S., Gollapalli, M., Ahmed, R., Aftab, S., Ahmad, M., Khan, M.A., & Mosavi, A. (2022). Rainfall prediction system using machine learning fusion for smart cities. Sensors, 22(9), 3504. https://doi.org/10.3390/s22093504
  31. Schulz, E., Speekenbrink, M., & Krause, A. (2018). A tutorial on Gaussian process regression: Modelling, exploring, and exploiting functions. Journal of Mathematical Psychology, 85, 1–16. https://doi.org/10.1016/j.jmp.2018.03.001
  32. Salahi, A., Ashrafzadeh, A., & Vazifedoust, M. (2024a). Assessing the forecasting accuracy of intense precipitation events in Iran using the WRF model. Earth Sciences Informatics, 17, 2199–2211. https://doi.org/10.1007/s12145-024-01274-x
  33. Salahi, A., Ashrafzadeh, A., & Vazifedoust, M. (2024b). Remote sensing-based precipitation forecasting using cloud optical characteristics: threshold optimization and evaluation in Northern and Western Iran. Natural Hazards, 120, 3661–3675. https://doi.org/10.1007/s11069-023-06352-9
  34. Sagi, O., & Rokach, L. (2018). Ensemble learning: A survey. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 8(4), e1249. https://doi.org/10.1002/widm.1249
  35. Tuysuzoglu, G., Birant, K.U., & Birant, D. (2023). Rainfall prediction using an ensemble machine learning model based on K-stars. Sustainability, 15(7), 5889. https://doi.org/10.3390/su15075889
  36. Taunk, K., De, S., Verma, S., & Swetapadma, A. (2019). A brief review of nearest neighbor algorithm for learning and classification. In 2019 International Conference on Intelligent Computing and Control Systems (ICCS) (pp. 1255–1260). IEEE. https://doi.org/10.1109/ICCS45141.2019.9065747
  37. Tukey, JW. (1977). Exploratory Data Analysis. https://doi.org/10.1007/978-3-031-20719-8_2
  38. Velasco, L.C., Aca-ac, J.M., Cajes, J.J., Lactuan, N.J., & Chit, S.C. (2022). Rainfall forecasting using support vector regression machines. International Journal of Advanced Computer Science and Applications, 13(3). https://doi.org/10.14569/IJACSA.2022.0130329
  39. Wang, Y., & Witten, I.H. (1997). Induction of model trees for predicting continuous classes. Proceedings of the Poster Papers of the European Conference on Machine Learning (pp. 48–53). https://hdl.handle.net/10289/1183
  40. Wu, Y., Wang, H., Zhang, B., & Du, K.L. (2012). Using radial basis function networks for function approximation and classification. ISRN Applied Mathematics. https://doi.org/10.5402/2012/324194
  41. Xu, J., & Zheng, Y. (2024). Remote sensing data analysis for urban planning and land use change. Environment Energy Earth Science, 3, 20–25. https://doi.org/10.62051/gwpk5m73

 

ارسال نظر در مورد این مقاله
نام را وارد کنید.
نشانی پست الکترونیکی را به درستی وارد کنید.
وابستگی سازمانی را به درستی وارد کنید.
توضیحات را وارد کنید (حداقل 50 حرف)
CAPTCHA Image
شناسه امنیتی را به درستی وارد کنید.
دوره 39، شماره 6 - شماره پیاپی 104
بهمن و اسفند 1404
صفحه 572-553

  • تاریخ دریافت 09 بهمن 1404
  • تاریخ بازنگری 07 خرداد 1405
  • تاریخ پذیرش 11 خرداد 1405
  • تاریخ اولین انتشار 11 خرداد 1405