Figure 2.
A diagram explains the transformer-based object detection model showing backbone, encoder, decoder, and prediction stages.The diagram presents the architecture of a transformer-based object detection model. The process starts with image input fed into a convolutional neural network backbone that extracts image features and adds positional encoding. The transformer encoder processes these features and sends them to the transformer decoder, which uses object queries. The prediction heads output classifications and bounding boxes for detected objects. The accompanying image displays two birds with bounding boxes around them, showing detected objects within the visual field.

DETR model architecture

Source(s): Figure courtesy of Carion et al. (2020) 

or Create an Account

Close subscription notice
Close access options