This study aims to apply a transformer-based network to achieve language-guided interactive segmentation for the automated analysis of ferrography images, while also alleviating the high cost of manual annotation.
To tackle the challenges of visual-linguistic alignment in referring image segmentation (RIS) and the complexity of ferrography images, a model named Multi-scale Fusion and Text-guided segmentation (MFT) is proposed. MFT injects textual cues into multi-scale visual features via a cross-modal fusion module. It then uses two sequential decoders to enhance cross-scale interaction and refine visual-text alignment for accurate segmentation. MFT is trained and evaluated on a wear debris data set with 11 categories. To reduce its annotation costs, a label enhancement strategy is introduced as a by-product. It leverages a pretrained segmentation model and multi-modal large language models (MLLMs) to automatically generate fine-grained masks and appearance-based descriptions from bounding box-annotated images, providing MFT mask-level supervision and rich textual guidance.
MFT achieves more accurate segmentation from referring expressions than other transformer-based methods. Moreover, MLLMs-generated descriptions guide the model more effectively than using category names alone.
MFT enables accurate interactive segmentation via multi-scale feature fusion and two sequential decoders, guided by either category names or appearance-based descriptions – especially effective with the latter. It achieves 72.20% mean Intersection over Union and 80.52% Prec@0.5 with only 124.6 M parameters, demonstrating competitive accuracy and efficiency. Combined with the enriched label, it also reduces annotation costs and improves segmentation quality.
The peer review history for this article is available at: https://publons.com/publon/10.1108/ILT-08-2025-0356/
