Article navigation
Purpose

The work presented in this study responds to the growing need for sustainable, human-centric and intelligent management of supply chain logistics in Industry 5.0. The work proposes the use of a reinforcement learning (RL) method in the management of inventories/storage and delivery in a way that also serves the purpose of the human decision-maker.

Design/methodology/approach

This work proposes a cyber-physical system (CPS)-based architecture incorporating a deep Q-network (DQN)-based RL model. The system dynamically interfaces with the environment of the supply chain to optimize its choices regarding routing, resource allocation and storage. The proposed framework models the decision problem in the logistics domain as a Markov Decision Process and uses a model-free deep reinforcement learning approach (deep Q-learning/DQN) to learn the policy in the decision process. A new multi-criteria approach has been described in the paper regarding the constraining objectives of logistics costs, service level requirements and efficiency considerations. The experiments are carried out using a large-scale database involving freight costs, order volume information, product distribution trends and transportation requirements.

Findings

The experimental results indicate that the new approach of RL reduces the costs of logistics by 15% and the service level of compliance by 20% relative to the existing method of optimization. The model has the ability to work perfectly under varying demanding conditions. The system also improves the effectiveness of the CPS through the enhancement of real-time response capabilities.

Research limitations/implications

Although the results clearly indicate the efficacy of the model regarding cost efficiency and performance improvement, there might be challenges involved in using large amounts of heterogeneous data in domains where the unavailability of quality data can be a problem. The generalization ability of the model when trained and tested using incomplete/noisy data in diverse industries can be investigated in the future. However, the usage of RL-based CPS models will definitely provide opportunities to businesses to achieve high levels of automation, robustness and sustainability in next-generation logistics.

Originality/value

This research has been novel in its merging of RL and cyber-physical system architecture in the development of a completely autonomous and adaptive decision-making model for the supply chain. Against the backdrop of conventional models that rely on static principles and/or heuristic development models that do not lean heavily toward continuous interaction between the environment and the learning model, the newly developed model of DQN-CPS takes the first step toward improving efficiency in global logistics through intelligent system development along the lines of Industry 5.0.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$41.00
Rental

or Create an Account

Close subscription notice
Close access options