Figure 4.
A context model combines spatial and temporal token mixing for frames t minus 2, t minus 1 and t, producing mean mu and scale sigma.The latent features from frame t minus 2 and frame t minus 1 are tokenised separately. Each token stream enters a local token mixer within spatial context modelling. The two local outputs are combined through an addition operation and passed to a joint token mixer within temporal context modelling. The latent features from frame t are tokenised and processed by masked self-attention. This output enters a global token mixer. The joint token mixer also feeds the global token mixer. The global output branches into a mean head and a scale head. The mean head produces mu, and the scale head produces sigma.

Architecture of the abstracted transformer entropy model (ATEM). The token mixer blocks are categorized into spatial and temporal context modeling modules. In this design, the left branch of each module processes latent features from previous frames (local token mixer for intra-frame tokens and joint token mixer for inter-frame tokens), while the right branch operates on the current frame, incorporating additional global token mixing layers to integrate historical context

Source: Authors’ own work

or Create an Account

Close subscription notice
Close access options