Underground inspection robots need reliable semantic perception in narrow, low-texture and unevenly illuminated spaces. This study aims to improve RGB-D semantic mapping for underground robotic inspection by generating more stable semantic labels and more coherent 3D semantic maps.
An attention-enhanced and boundary-refined DeepLabv3+ network is used as the semantic front end. A multi-head self-attention module is introduced to strengthen global contextual modeling in degraded regions, and a boundary-enhanced conditional random field is used to refine object contours. The pixel-level semantic labels are then projected into RGB-D point clouds through back-projection, multi-frame semantic fusion and semantic-geometric optimization. The method is evaluated on data collected by a wheeled-legged inspection robot equipped with a RealSense D455 camera in a simulated underground tunnel.
The proposed method improves mean intersection over union (mIoU) from 73.7% to 78.2% and mean pixel accuracy (mPA) from 86.5% to 92.1%. The 3D semantic mapping results remain relatively coherent under normal, high-illumination and low-illumination conditions. The mapping time increases after semantic fusion but remains within an acceptable range for offline inspection-map construction.
The method provides a semantic mapping solution for robotic inspection in underground spaces where visual degradation affects scene understanding, target localization and subsequent equipment inspection.
This study combines global semantic enhancement, boundary-aware refinement and RGB-D semantic projection in an underground robotic inspection framework. It links front-end semantic perception with 3D map construction rather than evaluating image segmentation alone.
