Figure 1.
A super-resolution framework shows shallow local feature encoding, D M Net processing, transformer blocks and pixel shuffle to produce I S R from I L R.The diagram shows shallow local feature encoding and reconstruction for super-resolution. I L R passes through convolution and splits into two branches labelled shallow local feature encoding, ours and shallow local feature encoding, others. In the upper branch, f passes through M p, D M Net and feature domain to produce Z 0 belongs to D 1. In the lower branch, f passes through others and feature domain to produce Z i belongs to D 2. A transformer D is shown with optimisation strategy defined as theta prime equals a r g min over theta in Theta of expectation over x h r, x s r in P L of double vertical line I S R minus I H R vertical line 1. In the main pipeline, I L R passes through convolution to channel feature map. Pixel mask M p produces I L R prime. D M Net with theta D and theta E processes features with n 36 arrow n 64 arrow ellipsis arrow n 1024 arrow n 512 arrow ellipsis arrow n 64. Features are combined by plus operations, then pass through transformer blocks and pixel shuffle, and produce I S R. The legend lists convolution, D W convolution, D W block, transformer, pixel shuffle, channel feature map, skip connection, and n 36 channel.

An overview of the proposed D-MNet transformer. Our method comprises shallow feature encoders, D-MNet and transformer-based decoders. The D-MNet components utilize depthwise convolution, which uses varying feature channels n (ranging from 32 to 1024) to achieve feature extraction. The encoded output Z0D1 is then used as input for the transformer decoder D. Mp represents the pixel-level masking strategy, ILR denotes the LR image, and ILR signifies the masked output. The B indicates the bicubic upsampling operation, and the final output ISR corresponds to the SR image. Best viewed in color

or Create an Account

Close Modal
Close Modal