Document Type : Original Article
Authors
Research Institute of Meteorology and Atmospheric Sciences (RIMAS), Tehran, Iran
10.22034/iwm.2026.2089458.1274
Abstract
Extended Abstract
Introduction: Soil moisture is a key state variable linking the land surface to the atmosphere and governing the water cycle, energy balance, and vegetation dynamics. Accurate and spatially continuous information on this variable is a prerequisite for sustainable water resources management, precision agriculture, flood forecasting, and drought monitoring, particularly in Iran, where recurrent and prolonged droughts and groundwater depletion make efficient water allocation an urgent priority. Ground-based measurements, although accurate, remain point-based and expensive and cannot portray the spatial heterogeneity of soil moisture at large scales, while national monitoring networks are still sparse. In recent decades, remote sensing technology and climate reanalysis products such as the Global Land Data Assimilation System (GLDAS) have emerged as powerful alternatives that provide continuous spatio-temporal fields of surface soil moisture. In parallel, machine learning algorithms, especially tree-based gradient boosting methods such as XGBoost, have shown an outstanding ability to capture complex nonlinear relationships among environmental variables. This study examines how far a data-driven model can reproduce the spatio-temporal behavior of surface soil moisture in Iran using only geographical coordinates and calendar time as predictors.
Materials and methods: Daily surface soil moisture from the GLDAS-CLSM V2.2 product (0.25° × 0.25° resolution) was acquired from NASA's GES DISC archive for the 21-year period 2005-2025, masked to Iran's political boundary (GADM), converted from kg m⁻² to equivalent water depth (mm), and aggregated to monthly means in the R environment. An XGBoost regression model was trained with four predictors only, namely longitude, latitude, month number, and year, which implicitly encode static environmental gradients (topography, soil, and climate) as well as seasonal and inter-annual signals. The dataset was randomly split into a training subset (80%) and an independent test subset (20%), and the hyperparameters (learning rate 0.05, maximum tree depth 8, subsampling ratio 0.8, and 300 boosting rounds) were tuned through 5-fold cross-validation on the training data. Model performance was evaluated with eight statistical metrics (R², NSE, KGE, Willmott's d, RMSE, MAE, Bias, and P-Bias), and seven standardized drought indices (SMI, SMA, SSI, SMD, RSMI, SMAP, and DSI) together with a percentile-based index (SMP) were derived from the predicted maps for drought monitoring.
Results and discussion: On the untouched test set, the model achieved R² = 0.966, NSE = 0.965, KGE = 0.960, d = 0.991, RMSE = 0.28 mm, and MAE = 0.18 mm, with bias close to zero and predictions tightly clustered around the 1:1 line. Monthly maps for 2020-2025 reproduced the expected spatial gradient, with the highest values along the Caspian coast and the northern slopes of Alborz and the lowest values in the arid central and south-eastern regions, together with a clear wet-winter/dry-summer cycle. Feature-importance analysis confirmed the dominance of the temporal variables.
Conclusion: With only four minimal inputs, XGBoost provides an accurate and computationally efficient tool for reconstructing and gap-filling GLDAS surface soil moisture at the national scale, which is valuable for drought early-warning systems, irrigation planning, and water resources management in ungauged regions. Future studies should apply spatial block or leave-one-year-out cross-validation to guarantee the independence of training and test data, and integrate higher-resolution satellite observations (e.g., Sentinel-1) and physical predictors such as precipitation, temperature, soil texture, and land cover to improve accuracy and transferability; the framework can also assess climate-change impacts on future soil moisture.
Keywords
Subjects