When materials are reflective, stacking is complex and hardware resources are limited, using machine vision for automated depalletizing becomes a challenging task. This paper aims to introduce a lightweight multimodal fusion semantic segmentation network, combined with point cloud data for three-degree-of-freedom pose estimation.
First, the proposed MIMF-IRA-Unet is an encoder–decoder architecture that takes input data from two different modalities, depth and intensity, obtained from a time-of-flight) camera. Second, two lightweight feature extraction modules are used to extract features from both modalities to enrich feature representation and reduce holes caused by reflections. Then, a multi-stage feature fusion process combines these features to enhance feature expression and better reconstruct the object. Finally, the prediction results are combined with the point cloud data to perform three-degree-of-freedom pose estimation.
MIMF-IRA-Unet achieved impressive results on a customized data set, with 90.16% intersection over union and 94.79% F1 scores with minimal parameters (1.2 M) and computational requirements (2.05 GFLOPs). Moreover, the short inference time (156 ms for two 320 × 240 images) not only outperforms the correlation baseline, but also maintains a good balance between accuracy and efficiency.
This study offers a solution for the automated picking of reflective materials in die-casting plants.
