نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسندگان English
Extended Abstract
Introduction
Landslides are among the most destructive geomorphological hazards worldwide, causing significant human casualties, infrastructure damage, and long-term socio-economic disruption. Their occurrence is governed by a complex interplay of natural and anthropogenic factors that destabilize slope systems. In recent decades, landslide susceptibility mapping (LSM) has evolved from qualitative and heuristic approaches toward quantitative, data-driven modeling frameworks supported by Geographic Information Systems (GIS) and Remote Sensing (RS). Within this paradigm, machine learning algorithms have emerged as powerful tools capable of capturing nonlinear and multidimensional relationships between landslide occurrences and environmental variables. Algorithms such as Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), and Gradient Tree Boosting (GTB) have demonstrated strong predictive performance in similar studies. Geomorphometric indices, including the Topographic Wetness Index (TWI), Stream Power Index (SPI), Normalized Difference Vegetation Index (NDVI), and curvature measures (plan, profile, and general curvature), play a critical role in explaining slope instability. However, the magnitude and interaction of these factors vary spatially, particularly in mountainous watersheds, leading to uncertainties in landslide prediction. The present study addresses this gap by evaluating the role of key geomorphic indicators and assessing the performance of four machine learning algorithms (RF, SVM, KNN, and GTB) in landslide hazard zonation within the Rezai-Chay watershed in Ardabil Province, Iran. The novelty of this research lies in the integrated application of multiple geomorphometric indices alongside advanced machine learning models in a relatively underexplored watershed.
Methodology
The study was conducted in the Rezai-Chay watershed, covering approximately 180 km², characterized by diverse topography, elevations ranging from 1133 to 2480 meters, and mixed land uses including rangelands, agriculture, and rural settlements. The methodological framework consisted of several key stages: data collection, preprocessing, feature extraction, model training, and validation. A comprehensive set of environmental variables was selected based on literature review, field observations, and data availability. These included elevation, slope, aspect, lithology, land use, rainfall, distance to rivers, faults, and roads, as well as geomorphometric indices such as TWI, SPI, curvature types, and NDVI. Spatial datasets were derived from multiple sources, including ASTER DEM (30 m resolution), Sentinel-2 imagery, GLC-FCS30D land cover data, geological maps, and meteorological records. All layers were processed and standardized to a 30 m spatial resolution and normalized to a 0–1 range. The landslide inventory map was prepared using official records, satellite image interpretation (Google Earth), and field verification. The dataset was divided into training (70%) and testing (30%) subsets. To ensure balanced classification, non-landslide samples were randomly generated equal in number to landslide points. Four machine learning models (RF, SVM, KNN, and GTB) were implemented in the Google Earth Engine environment. Model performance was evaluated using Receiver Operating Characteristic (ROC) curves and Area Under the Curve (AUC), along with overall accuracy. These metrics provided insights into the predictive capability and classification reliability of each model.
Results and discussion
The results indicate that all four models performed exceptionally well, with AUC values ranging from 0.979 to 0.984, confirming their high predictive capability. Among them, the RF model achieved the highest performance (AUC = 0.984), followed closely by KNN (0.982), GTB (0.981), and SVM (0.979). Despite similar overall accuracy, significant differences were observed in the spatial distribution of hazard classes. The RF model produced the most balanced classification, allocating approximately 21.66% of the area to very high risk and maintaining a realistic distribution across all classes. In contrast, KNN overestimated high-risk zones (25.85%), likely due to its sensitivity to local data patterns and noise. SVM exhibited a conservative tendency, underestimating high-risk areas (16.61%), while GTB showed performance close to RF but with a slight bias toward lower-risk classes. From a geomorphological perspective, slope, distance to rivers, and TWI were identified as the most influential factors. Areas with slopes greater than 30° and proximity less than 300 meters to streams exhibited the highest susceptibility. High TWI values (>8), indicating soil moisture accumulation, significantly contributed to slope instability, especially when combined with steep gradients. Curvature analysis revealed that concave (negative curvature) zones, associated with water convergence and increased pore pressure, corresponded strongly with high-risk areas, particularly in RF and GTB outputs. The SPI index showed stronger influence in KNN results, reflecting its sensitivity to localized erosion processes. NDVI analysis indicated that areas with sparse vegetation cover were more prone to landslides, although its importance varied across models. Conversely, factors such as distance to faults, roads, and slope aspect had limited influence in this watershed, suggesting that geomorphic and hydrological controls dominate landslide occurrence in the study area.
Conclusion
This study demonstrates the effectiveness of machine learning algorithms in landslide susceptibility mapping within complex mountainous environments. While all models exhibited high predictive accuracy, the Random Forest algorithm outperformed others due to its balanced classification, robustness against overfitting, and strong alignment with geomorphological processes. The GTB model also showed reliable performance and can be considered a suitable alternative. The findings highlight that model selection significantly influences hazard zonation outcomes. KNN tends to overestimate, while SVM may underestimate risk levels, limiting their standalone application in decision-making. Therefore, ensemble approaches or hybrid models are recommended for future studies to enhance prediction reliability. Additionally, the study confirms that slope gradient, hydrological proximity, and soil moisture conditions are the primary drivers of landslide occurrence in the Rezai-Chay watershed. Incorporating higher-resolution spatial data and temporal environmental variability in future research could further improve model performance and hazard assessment accuracy. Ultimately, the results provide a scientific basis for risk management, land-use planning, and mitigation strategies in landslide-prone regions.
Keywords: Landslide, Geomorphology, Receiver Operating Characteristic (ROC) Curve, Machine Learning, Rezi Chay.
کلیدواژهها English