Article navigation
Purpose

When materials are reflective, stacking is complex and hardware resources are limited, using machine vision for automated depalletizing becomes a challenging task. This paper aims to introduce a lightweight multimodal fusion semantic segmentation network, combined with point cloud data for three-degree-of-freedom pose estimation.

Design/methodology/approach

First, the proposed MIMF-IRA-Unet is an encoder–decoder architecture that takes input data from two different modalities, depth and intensity, obtained from a time-of-flight) camera. Second, two lightweight feature extraction modules are used to extract features from both modalities to enrich feature representation and reduce holes caused by reflections. Then, a multi-stage feature fusion process combines these features to enhance feature expression and better reconstruct the object. Finally, the prediction results are combined with the point cloud data to perform three-degree-of-freedom pose estimation.

Findings

MIMF-IRA-Unet achieved impressive results on a customized data set, with 90.16% intersection over union and 94.79% F1 scores with minimal parameters (1.2 M) and computational requirements (2.05 GFLOPs). Moreover, the short inference time (156 ms for two 320 × 240 images) not only outperforms the correlation baseline, but also maintains a good balance between accuracy and efficiency.

Originality/value

This study offers a solution for the automated picking of reflective materials in die-casting plants.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$39.00
Rental

or Create an Account

Close subscription notice
Close access options