Article navigation

The selection of negative samples is critical to landslide susceptibility assessment, yet the effect of their spatial distance from landslide locations remains insufficiently understood. This study investigated how buffer-zone sampling ranges affect the generalisation performance of random forest models in Lin’an District, Hangzhou, China. Using a 1:1 ratio of positive to negative samples, six datasets were constructed with negative samples randomly selected within buffer ranges of 0.5–1.5 to 0.5–6 km. Nine conditioning factors were used for model development, and SHAP (SHapley Additive exPlanations) analysis was employed for model interpretation. Results showed a clear scale-dependent effect of buffer distance. Model performance improved with increasing sampling range up to an optimal limit, with the 0.5–5 km buffer achieving the highest predictive accuracy and most reliable susceptibility zonation. Further expansion to 0.5–6 km reduced model performance. SHAP analysis indicated that sampling scale altered the statistical representation of negative samples, thereby affecting feature importance and model responses. These findings demonstrate that buffer distance should be evaluated as a sampling-scale parameter rather than selected solely by empirical convention, providing a practical basis for optimising negative-sample selection in regional landslide susceptibility modelling.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$39.00
Rental

or Create an Account

Close subscription notice
Close access options