The diagram presents the architecture of a transformer-based object detection model. The process starts with image input fed into a convolutional neural network backbone that extracts image features and adds positional encoding. The transformer encoder processes these features and sends them to the transformer decoder, which uses object queries. The prediction heads output classifications and bounding boxes for detected objects. The accompanying image displays two birds with bounding boxes around them, showing detected objects within the visual field.DETR model architecture
Source(s): Figure courtesy of Carion et al. (2020)
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.