Article navigation
Purpose

The purpose of this paper is to study a reinforcement learning training method with action decoupling and experience transfer capabilities, aiming to overcome the difficulty of robots learning composite actions caused by sparse rewards. When learning complex, compound actions under sparse reward environments, robots require an enormous number of training steps to trigger an effective reward by chance, which can indefinitely prolong the training cycle or even lead to learning failure. Therefore, this paper proposed a method to alleviate the sparsity issue through action decomposition and experience aid.

Design/methodology/approach

A reinforcement learning method, termed experience-acquirer-aided training (EAAT), is proposed to enhance the reinforcement learning capability of robots when learning complex actions under sparse rewards. EAAT integrates two identical actor-critic reinforcement learning algorithms, which focus on local and global compound actions respectively with differential observations. Through mutual fusion via time-variant Q-functions and their respective critic networks – where local critics act as experience acquirers and continuous reward generators – this method enables efficient learning in sparse reward scenarios. Moreover, a fuzzy reward function is designed based on Dempster–Shafer fusion theory, which reduces the complexity of designing traditional continuous reward functions and provides further efficiency support for EAAT.

Findings

Core experimental validation on a robotic autonomous precision assembly platform demonstrates that EAAT achieves high-quality convergence with 50% fewer required time steps. In addition, the approach effectively generalizes to inverse kinematics learning, thereby broadening its application scope.

Originality/value

This work introduces a novel hybrid training method EAAT, a non-hierarchical dual-agent parallel architecture that can automate reward generation via a pretrained critic network and integrate fuzzy evaluation to simplify reward design. In addition, for specific peg-hole assembly tasks, an innovative hole-searching strategy and improved insertion control methods are added.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$39.00
Rental

or Create an Account

Close subscription notice
Close access options