The diagram shows shallow local feature encoding and reconstruction for super-resolution. I L R passes through convolution and splits into two branches labelled shallow local feature encoding, ours and shallow local feature encoding, others. In the upper branch, f passes through M p, D M Net and feature domain to produce Z 0 belongs to D 1. In the lower branch, f passes through others and feature domain to produce Z i belongs to D 2. A transformer D is shown with optimisation strategy defined as theta prime equals a r g min over theta in Theta of expectation over x h r, x s r in P L of double vertical line I S R minus I H R vertical line 1. In the main pipeline, I L R passes through convolution to channel feature map. Pixel mask M p produces I L R prime. D M Net with theta D and theta E processes features with n 36 arrow n 64 arrow ellipsis arrow n 1024 arrow n 512 arrow ellipsis arrow n 64. Features are combined by plus operations, then pass through transformer blocks and pixel shuffle, and produce I S R. The legend lists convolution, D W convolution, D W block, transformer, pixel shuffle, channel feature map, skip connection, and n 36 channel.An overview of the proposed D-MNet transformer. Our method comprises shallow feature encoders, D-MNet and transformer-based decoders. The D-MNet components utilize depthwise convolution, which uses varying feature channels n (ranging from 32 to 1024) to achieve feature extraction. The encoded output is then used as input for the transformer decoder D. represents the pixel-level masking strategy, denotes the LR image, and signifies the masked output. The indicates the bicubic upsampling operation, and the final output corresponds to the SR image. Best viewed in color
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.