Castillo-Martínez et al. (2026) Assessing the Effect of Training Record Length on Daily Pan Evaporation Estimation Using MLR, MLP, LSTM, and XGBoost Models in a Semi-Arid Region of Mexico
⚠️ Warning: This summary was generated from the abstract only, as the full text was not available.
Identification
- Journal: Water
- Year: 2026
- Date: 2026-09-25
- Authors: Luis Fernando Castillo-Martínez, Luis Octavio Solís-Sánchez, Mireya Moreno-Lucio, Verónica Libertad Medina-Llamas, Celina Lizeth Castañeda-Miranda, Carlos Alberto Olvera-Olvera, Ramón Jaramillo-Martínez, José Israel Casas-Flores
- DOI: 10.3390/w18192392
Research Groups
Not explicitly stated in the provided text. The study was conducted in a semi-arid region of Mexico.
Short Summary
This study evaluated the impact of training record length (10 vs. 20 years) on the performance of various machine learning models for daily pan evaporation estimation in a semi-arid region, finding that increased record length does not guarantee uniform performance improvements across all models.
Objective
- To assess how the length of the training record affects the accuracy of Multiple Linear Regression (MLR), Multilayer Perceptron (MLP), Long Short-Term Memory (LSTM), and Extreme Gradient Boosting (XGBoost) models in estimating daily pan evaporation.
Study Configuration
- Spatial Scale: A semi-arid region of Mexico.
- Temporal Scale: Daily observations over two historical data scenarios: 10 years (2015–2022) and 20 years (2005–2022), with an independent testing period (2023–2024).
Methodology and Data
- Models used: Multiple Linear Regression (MLR), Multilayer Perceptron (MLP), Long Short-Term Memory (LSTM), Extreme Gradient Boosting (XGBoost).
- Data sources: Historical daily observations of temperature, relative humidity, wind speed, and solar radiation.
Main Results
- For the 10-year training scenario, MLR - M5 achieved the most favorable overall performance with a Root Mean Square Error (RMSE) of 1.98 mm/day.
- For the 20-year training scenario, LSTM - M9 was selected as the reference configuration, achieving an RMSE of 1.97 mm/day.
- Extending the training record from 10 to 20 years reduced LSTM RMSE by 0.12 mm/day (5.74%).
- Conversely, extending the training record from 10 to 20 years increased MLR RMSE by 0.05 mm/day (2.53%).
- The 20-year LSTM RMSE (1.97 mm/day) was only 0.01 mm/day (0.51%) lower than the 10-year MLR RMSE (1.98 mm/day).
- The study concludes that increasing the training record length does not guarantee uniform improvements and its effect depends on the specific learning paradigm and estimation configuration.
Contributions
- Quantifies the differential impact of training record length on various machine learning models (MLR, MLP, LSTM, XGBoost) for daily pan evaporation estimation.
- Highlights that longer training records do not universally improve model performance, with some models benefiting (e.g., LSTM) while others may degrade (e.g., MLR).
- Provides specific performance metrics (RMSE) for different model configurations and training data lengths in a semi-arid context.
Funding
Not explicitly stated in the provided text.
Citation
@article{CastilloMartínez2026Assessing,
author = {Castillo-Martínez, Luis Fernando and Solís-Sánchez, Luis Octavio and Moreno-Lucio, Mireya and Medina-Llamas, Verónica Libertad and Castañeda-Miranda, Celina Lizeth and Olvera-Olvera, Carlos Alberto and Jaramillo-Martínez, Ramón and Casas-Flores, José Israel},
title = {Assessing the Effect of Training Record Length on Daily Pan Evaporation Estimation Using MLR, MLP, LSTM, and XGBoost Models in a Semi-Arid Region of Mexico},
journal = {Water},
year = {2026},
doi = {10.3390/w18192392},
url = {https://doi.org/10.3390/w18192392}
}
Original Source: https://doi.org/10.3390/w18192392