Masonry construction is well suited to robotic automation because it involves the sequential assembly of standardized blocks. However, existing robotic masonry systems depend heavily on manual programming. This study develops a curiosity-driven reinforcement learning (RL) framework that allows robots to learn control policies through environmental interaction rather than task-specific programming, with a focus on dry-stack masonry block assembly within a physics-based simulation environment.
A curiosity-driven RL framework is proposed by integrating an Intrinsic Curiosity Module (ICM) into the Proximal Policy Optimization (PPO) algorithm. Intrinsic novelty-based rewards are combined with extrinsic task rewards to enhance exploration and learning efficiency. A physics-based simulation environment is developed using PyBullet. Multiple dry-stack masonry wall configurations are evaluated under dense and sparse reward conditions.
For simple tasks with dense rewards, PPO and PPO + ICM exhibit comparable performance. As task complexity increases, PPO + ICM consistently outperforms PPO; for the most demanding 4 × 3 Running Wall, it achieves up to 20.2% higher average rewards and 27.7% higher average success rates. Under sparse reward conditions, PPO fails to learn effective policies, whereas PPO + ICM maintains stable learning and high task success.
The simulation employs a restricted action space limited to joint rotations and does not model collision dynamics. Future work should incorporate extended action spaces, collision effects and physical-robot validation to assess real-world applicability.
The proposed method reduces reliance on task-specific programming by enabling robots to adapt to varying block arrangements and construction sequences. This supports more flexible deployment of robotic systems for masonry construction.
This study presents a novel application of curiosity-driven RL to long-horizon dry-stack masonry block assembly in simulation. It offers a transferable paradigm for addressing sparse rewards and long-horizon decision-making in construction robotics.
