The input frame x sub t enters the context encoder and the motion-estimation branch. The context encoder produces latent representation y sub t, which passes through Q, A E and arithmetic encoding. The resulting bitstream is processed by arithmetic decoding to obtain y hat sub t. The context decoder uses y hat sub t and motion context to produce reconstructed frame x hat sub t. In the motion branch, motion estimation produces v sub t. The motion encoder transforms v sub t before Q, A E and arithmetic encoding. The motion entropy model receives encoded motion information and provides distribution parameters sigma and mu to the arithmetic encoder and arithmetic decoder. Arithmetic decoding produces v hat sub t, which enters the motion decoder. The decoded motion output feeds the motion-context module. Motion context is supplied to the context encoder, the context decoder and the upper entropy model. The upper entropy model also provides sigma and mu to its arithmetic encoder and arithmetic decoder. Reconstructed frame x hat sub t is fed back for subsequent motion estimation.Illustrations of a generalized design of motion compensation-based LVC architectures. Motion context-based methods, exemplified by the DCVC series (Li et al., 2021, 2022; Sheng et al., 2021), typically employ two parallel entropy coding pipelines: one for motion context mining and another for latent feature coding, with the former serving as prior knowledge for the latter. In this figure: and are the original and reconstructed frames at time stamp . is latent feature encoded by the context encoder and is reconstructed latent feature decoded from the bitstream by an arithmetic decoder . stands for the quantization process, is the arithmetic encoder. and are means and scales predicted by the entropy model for arithmetic coding and decoding
Source: Authors’ own work
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.