Article navigation

This study addresses the route stowage planning problem in inland container shipping using a multi-stage stochastic programming model (SPM). The model dynamically optimises stack occupancy and deviations between adjacent stages’ stowage plans by way of rolling scheduling, explicitly integrating dynamic uncertainties like stochastic container volume variations and specific seasonal waterway constraints. A robust optimisation approach with interval estimation converts the SPM into a mixed-integer programming model (MIPM) for mathematical solver accessibility. An adaptive reinforcement learning framework based on proximal policy optimisation (PPO) is proposed, featuring enhanced policy updates and adaptive exploration. Computational results show that both MIPM and PPO outperform Deep Q Network algorithms across scales. For large-scale problems, PPO achieves solutions within 30 s on average, matching or surpassing MIPM in efficiency. PPO maintains stability under ≤10% demand perturbations but degrades at 15% due to complexity. Hyperparameter analysis confirms the balanced configurations optimise exploration trade-offs, ensuring convergence reliability.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$41.00
Rental

or Create an Account

Close Modal
Close Modal