Event camera sensors are bio-inspired sensors that asynchronously capture per-pixel brightness changes and output a stream of events encoding the polarity, location and time of these changes. These systems are witnessing rapid advancements as an emerging field, driven by their low latency, reduced power consumption and ultra-high capture rates. This survey explores the evolution of fusing event stream captured with traditional frame-based capture, highlighting how this synergy significantly benefits various video restoration and 3D reconstruction tasks. This paper systematically reviews major deep learning contributions that use unique data from event camera systems for image/video enhancement and restoration, focusing on temporal improvements, such as frame interpolation and motion deblurring, as well as spatial enhancements, including super-resolution, low-light/high dynamic range processing and artifact reduction. This paper also explores how the 3D reconstruction domain evolves with the advancement of event-driven fusion. Diverse topics are covered, with in-depth discussions on recent works for improving visual quality under challenging conditions. Additionally, the survey compiles a comprehensive list of openly available data sets, enabling reproducible research and benchmarking. By consolidating recent progress and insights, this survey aims to inspire further research into leveraging event camera systems, especially in combination with deep learning, for advanced visual media restoration and enhancement.

Vision is the dominant human sense for perceiving the world and, in combination with the brain, facilitates learning. Traditionally, frame-based Red Green Blue (RGB) cameras have been the preferred choice for visual sensing, as they capture rich color, texture and semantic information. However, they struggle in challenging conditions such as extreme lighting and fast relative motion. Event cameras, also known as dynamic vision sensors (Brandli et al., 2014; Gallego et al., 2020), operating on principles similar to the human retina, offer a novel approach to capturing changes in a visual scene. Inspired by biological vision, event cameras capture asynchronous pixel-wise intensity changes, offering advantages such as low latency, high speed and high dynamic range (HDR).

Event camera systems offer unique advantages over traditional frame-based cameras through their novel approach to capturing visual information (Gallego et al., 2020). Unlike conventional frame-based cameras that record absolute intensity at fixed intervals, event cameras asynchronously detect per-pixel logarithmic brightness changes (⁠ΔL/Δt⁠), generating sparse spatiotemporal event streams with microsecond temporal resolution (⁠>10,000 fps) and HDR (⁠>120dB⁠). These attributes uniquely position them to address critical limitations in traditional imaging systems – motion blur in high-speed scenarios, over/under exposure in extreme lighting and bandwidth inefficiencies in redundant static scenes. By capturing the changes in the scene asynchronously, they allow for high temporal resolution and reduced motion blur, making them particularly suitable for image and video restoration applications under extreme capture conditions.

Since the event camera sensor only captures the changes in brightness asynchronously, it concurrently loses the grayscale information of the scene. Traditionally, the grayscale information is captured by an active pixel sensor, which is co-located with the event sensor in the same chip (Brandli et al., 2014). Downstream tasks use event integration, i.e. the fusion of the event data stream with grayscale capture for image reconstruction. Bao et al. (2024) introduce a novel method for intensity image creation using the information about the time of event emission for each pixel. Lei et al. (2024) study the impact of the number of events needed for image reconstruction using event cameras. They study fixed duration and fixed number of events as their primary reconstruction strategies and provide insights into the tradeoff between temporal resolution and image quality for each strategy.

Recent advancements in deep learning have significantly accelerated event camera applications by developing architectures tailored to their asynchronous, sparse data streams. Transformers (Vaswani et al., 2017) and spiking neural networks (SNNs) (Tavanaei et al., 2019) can process event-based spatiotemporal features directly, enabling tasks like high-speed video reconstruction (Liu et al., 2025b) and low-light enhancement (Liu et al., 2023). For instance, hybrid SNN-convolutional neural networks (CNN) pipelines (Lee et al., 2020) use spiking layers to extract microsecond-level temporal cues, fused with CNNs for spatial super-resolution, achieving better optical flow estimation than frame-based methods. Physics-informed networks further integrate event generation models (e.g. ΔL=∇L·v+∂L/∂t⁠) as differentiable layers, improving optical flow estimation accuracy in extreme lighting. Attention mechanisms, such as cross-modal asymmetric transformers, dynamically weight event and RGB features to reduce motion blur artifacts in deblurring tasks (Sun et al., 2022b). These methods leverage event cameras’ μ-second-resolution temporal data to overcome traditional vision bottlenecks, making them highly beneficial for applications in autonomous navigation, computational photography and real-time augmented reality. Deep learning can unlock unprecedented visual fidelity in dynamic environments by aligning architectural innovations with event cameras’ innate strengths – low latency, HDR and motion robustness.

Motivation. In this survey paper, we focus on the recent developments in using event cameras for image and video restoration applications. There exist excellent survey papers (Chakravarthi et al., 2024; Gallego et al., 2020; Shariff et al., 2024; Zheng et al., 2023a) covering the working principles and broad applications of event camera systems. Compared to them, in this survey, we intend to focus on the restoration and enhancement capabilities offered by event camera sensors when used in conjunction with traditional frame-based cameras. We explore different image and video capture challenges, such as low-light enhancement, HDR reconstruction, artifact reduction and focus control, as well as how it unlocks new capabilities, such as 3D reconstruction.

Key contributions. This survey makes several key contributions to the field of event-based vision. First, we provide a holistic and structured review of the synergy between event cameras and traditional frame-based sensors, bridging two often separate areas of research. Second, this work systematically categorizes major deep learning advancements into distinct temporal and spatial enhancement dimensions, offering a clear taxonomy for understanding the state-of-the-art in video restoration. Third, we extend our analysis beyond 2D restoration to the domain of 3D reconstruction, exploring how event streams are advancing tasks in this challenging area. Finally, by compiling a comprehensive list of publicly available data sets, this survey serves as a valuable, reproducible resource for both newcomers and seasoned researchers, aiming to standardize benchmarking and inspire future innovation in the field.

Outline. We start the survey by presenting a primer on event camera systems. After discussing the working principles, we discuss event data simulators and event data enhancement methods to round up the data acquisition task. In the following two sections, we present how event cameras can be leveraged to enhance visual content in spatial and temporal domains. Next, we discuss how event camera data can be leveraged for 3D scene reconstruction and survey recent work leveraging event data for the same. Finally, we include a comprehensive list of publicly available event camera data sets with the hope that curious readers may find it interesting to experiment with event data streams. We provide a map of this survey in Figure 1.

Figure 1.
A flowchart outlining the organization of this study on event camera guided visual media restoration and 3D reconstruction.The flowchart presents a structured framework for event camera guided visual media restoration and 3 D reconstruction. The top level introduces the overall theme and branches into 4 main sections. The first section describes event camera systems and data processing including event camera theory, event data representation, event camera simulator, and event data enhancement with components such as event data training, efficient event feature extraction, spatial and temporal up conversion, event denoising, and multi modal security processing. The second section focuses on temporal restoration including video reconstruction using handcrafted priors, model based approaches, C N N and R N N based approaches, followed by video interpolation and motion deblurring. The third section addresses spatial restoration including super resolution, low light enhancement, rain removal, and focus control for automatic focusing and all in focus imaging. The fourth section covers 3 D reconstruction including N e R F models with events, 3 D G S models with events, H D R conversion, and occlusion removal. The structure progresses hierarchically from foundational event camera technology to advanced restoration and reconstruction tasks.

Organization of the study

Figure 1.
A flowchart outlining the organization of this study on event camera guided visual media restoration and 3D reconstruction.The flowchart presents a structured framework for event camera guided visual media restoration and 3 D reconstruction. The top level introduces the overall theme and branches into 4 main sections. The first section describes event camera systems and data processing including event camera theory, event data representation, event camera simulator, and event data enhancement with components such as event data training, efficient event feature extraction, spatial and temporal up conversion, event denoising, and multi modal security processing. The second section focuses on temporal restoration including video reconstruction using handcrafted priors, model based approaches, C N N and R N N based approaches, followed by video interpolation and motion deblurring. The third section addresses spatial restoration including super resolution, low light enhancement, rain removal, and focus control for automatic focusing and all in focus imaging. The fourth section covers 3 D reconstruction including N e R F models with events, 3 D G S models with events, H D R conversion, and occlusion removal. The structure progresses hierarchically from foundational event camera technology to advanced restoration and reconstruction tasks.

Organization of the study

Close Figure 1.

Event cameras are bio-inspired vision sensors that capture visual information as asynchronous events triggered by changes in brightness. Different from traditional cameras, which capture images at fixed sampling intervals (e.g. 30 fps), event camera systems output a continuous stream of events, each containing a pixel address, timestamp and polarity (increase or decrease) of the brightness change.

Each pixel in an event camera operates independently, continuously comparing its current brightness level to a stored reference level. If the difference exceeds a predefined threshold, the pixel generates an event and resets its reference level to the current brightness. This process allows event cameras to capture motion and changes in the scene with very low latency (on the order of microseconds) and a HDR (typically 120 dB).

The working principle of an event camera at a circuit level can be summarized as follows:

  • Incoming light is processed by each pixel independently, continuously and asynchronously.

  • The photodiodes convert the incoming light into an electrical current, which is subsequently transformed into a corresponding voltage signal.

  • This voltage signal is compared to a reference voltage on a logarithmic scale to detect any changes in light intensity.

Formally, each event can be interpreted as a tuple (x,y,p,t)⁠, where (x,y) represents the pixel location, t the timestamp and p∈{−1,+1} is the polarity indicating the direction of brightness change. An event is triggered whenever a change in the logarithmic intensity L surpasses a predefined threshold C. This can be represented as follows:

(1)

where Δt is the time interval since the last event at pixel (x,y)⁠. We can represent the change in brightness since the last event at a pixel by ΔL(x,y,t)=L(x,y,t)−L(x,y,t−Δt), and hence, the ΔL(x,y,t)=p·C⁠. Furthermore, a stream of events can be represented as follows:

(2)

An example image of event camera capture is given in Figure 2. In Figure 2(a), a frame captured by an RGB camera is shown where the camera is moving from left to right and the cyclist is moving from right to left. Figure 2(b) shows the events captured over a small interval of time from the same scene. In the event data stream, red represents the regions where intensity decreased during capture and blue represents the regions where intensity increased. As we can see, in the parked e-scooter, as the white frame of the e-scooter moves from right to left, blue dots show up on the front side of the e-scooter, followed by red dots on the backside of the main axle, indicating how the high-intensity white region moved across in the frame. Similarly, as the black t-shirt cyclist moves fast from left to right, we see trailing blue dots signifying an increase in intensity due to the dark color of the t-shirt covering the regions and leading blue dots signifying bright skin covering the regions. For the purpose of visualization, we have projected the event data stream across time into a 2D plane, where the changes are visible. We provide a detailed comparison of a traditional RGB camera with an event camera in Table 1.

Figure 2.

An example of a captured intensity frame and its corresponding events between this captured intensity frame and the next frame. Red and blue dots in the event frame represent negative and positive polarity, respectively. The continuous event data stream is projected into a 2D frame for visualization

Figure 2.

An example of a captured intensity frame and its corresponding events between this captured intensity frame and the next frame. Red and blue dots in the event frame represent negative and positive polarity, respectively. The continuous event data stream is projected into a 2D frame for visualization

Close Figure 2.
Table 1.

Comparison of traditional RGB camera with event camera

CharacteristicsRGB cameraEvent camera
Capture methodFixed rate full frameAsynchronous per-pixel
Temporal resolutionMillisecondsMicroseconds
LatencyHigh, due to full frame captureLow, due to asynchronous events
Data outputFixed-frame rate with absolute brightness valuesSparse data stream encoding brightness changes
Dynamic rangeModerate, around 60 dBVery high, around 140 dB
Power consumptionHigh, due to continuous processingLow, due to sparse processing
ApplicationsGeneral purpose imaging, videoHigh-speed vision, HDR imaging

Event camera systems focus on changes in light intensity rather than absolute light levels and hence, only relevant information regarding a scene is captured, thus minimizing redundancy. Furthermore, since the event camera detects changes in logarithmic scale rather than absolute values, common issues like overexposure and underexposure, prevalent in traditional systems, can be avoided.

The sparse nature of the event data stream [equation (2)] makes it difficult to apply the deep neural network models predominantly designed for frame-based cameras. This necessitates the development of alternative representation formats for event data that can capture the visual information power appropriately. Some of the popular event representation methods (Zheng et al., 2023a) are described below.

Image based representation. Stacking asynchronous events into a synchronous 2D image representation, similar to frame-based cameras, is a straightforward solution to adapt events to existing deep learning (DL) methods. It should be noted that these representations are often set to preserve polarities, timestamps and event counts. Popular stacking approaches are:

  • Based on polarity, two separate channels can be used to evaluate the histogram of positive and negative events. The output of the event camera is collected into frames over a specified time interval T, using a separate channel for the event polarity. This can be finally fused into synchronous time events (Maqueda et al., 2018).

  • Based on timestamps, the events can be aggregated into synchronous frames to capture holistic information (Bai et al., 2022; Deng et al., 2020; Wang et al., 2019). The time duration of the event stream is divided into n equal-scale portions and the n grayscale frames are formed by merging the events in each interval through pointwise summation. These grayscale frames are further stacked together and are presented as input to the model.

  • Based on the number of events (Hu et al., 2020; Messikommer et al., 2020), propose a strategy to sample and stack events in a fixed constant number. By only updating the locations where new events are recorded, this method, when used in conjunction with sparse convolutions, can also maintain the temporal sparsity of events.

Surface based representation. Surface representation aims to map the event streams to a time-dependent surface and tracks the activity around the spatial location of the latest event. Usually, time t is represented as a monotonically increasing function of position (x,y)⁠. Surface of active events (Benosman et al., 2013) captures a time surface that encodes the time context in a neighborhood region of the event and helps to maintain spatio-temporal information for downstream tasks. Because of the monotonically increasing nature of timestamps, different normalization methods are developed for temporal invariant data representations (Afshar et al., 2019; Alzugaray and Chli, 2018; Lagorce et al., 2016).

Voxel based representation. Voxel representation maps the raw events into the nearest temporal grid within temporal bins. By partitioning both space and time into discrete levels, voxel-based representation quantizes and accumulates the events into voxel grids in (x,y,t) space. A pioneering work in this direction, Zihao Zhu et al. (2018) propose to insert events into volumes using a linearly weighted accumulation to improve the resolution along the temporal domain. To compactly capture raw information temporal spikes with minimal information loss, Baldwin et al. (2022) propose the concept of time-ordered recent event volume. The graph is constructed by voxelizing the spatio-temporal event, subsampling voxels to get representative voxels as vertice and finally integrating internal events in a voxel along the time axis to obtain the node features. The proposed EV-VGCNN is capable of finding neighbors and calculating edge weights for vertices. Choudhury et al. (2025) proposed a novel triplane and probabilistic autoencoder framework for compact and unified event stream representation, featuring a two-stage training scheme and Poisson-based voxel encoding that enables efficient reconstruction and compatibility with diffusion models.

Graph based representation. Graph representation transforms the sparse events within a time window into a set of connected nodes. For the problem of object detection, Bi et al. (2020) propose a residual graph CNN architecture to obtain a compact graph representation. Each node is a sampled event from a time-space voxel grid and nodes are connected with edges only if they have a weighted spatio-temporal Euclidean distance. Deng et al. (2022) propose a lightweight voxel graph CNN, which has been shown to achieve high accuracy with low model complexity.

Spike based representation. Spike representation uses SNNs to extract features from event streams asynchronously to solve diverse tasks (Gu et al., 2020; Orchard et al., 2015). However, complex dynamics and the non-differentiable nature of SNNs severely limit the applicability of these methods to large-scale problems.

Learning based representation. Learning based representation leverages the power of data to learn optimal representations to convert an asynchronous event stream to flexible representations that can be used for diverse tasks. Popular methods use multi-layer perceptron (MLP) architectures to aggregate spatio-temporal information (Gehrig et al., 2019), long-short term memory (LSTM) (Cannici et al., 2020) to integrate information into the temporal axis. Because it is fully differentiable, this method allows for the extraction of the most relevant representation for downstream tasks from the data itself. Huang et al. (2025a) present a learning-based framework for efficiently representing event voxel grids, which are typically sparse and inefficient to store due to the spatiotemporal nature of event camera data. Three neural representations – MLP, tensor decomposition and hash encoding – along with their respective strengths and limitations in enabling deep learning models to process event data for vision tasks. Choudhury and Su (2025) introduce a modified neural field model, composed solely of a MLP, to effectively represent sequences of event frames generated by event cameras, where each pixel encodes the polarity of brightness change within a temporal window. The more advanced learning based event feature extraction techniques and architectures are discussed in Section 2.4.1.

While event camera systems offer unparalleled advantages in terms of HDR, no motion blur and asynchronous sensing, their scarcity and high non-production-scale hardware cost remain challenging for the research community when acquiring event camera data streams. Hence, obtaining synthetic data for exploratory research and algorithm validation in a controlled and cost-efficient setting is necessary. This has led to the development of several event camera simulator systems that can generate large amounts of affordable and reliable event data.

Based on the underlying principle of event simulation, modern event simulators are broadly classified into three major groups:

Optimization based simulators. Optimization-based simulators operate based on optimizing hand-crafted modules for generating event data (Hu et al., 2021; Palinauskas et al., 2023; Rebecq et al., 2019; Ziegler et al., 2023). ESIM (Rebecq et al., 2019) can generate events, standard images and inertial measurements while simulating arbitrary camera motion in 3D scenes. It also provides complete ground truth data, including camera pose, velocity, depth maps and optical flow maps. However, ESIM does not perform noise modeling and is limited to synthetic 3D scenes or high-framerate input video. V2E (Hu et al., 2021) generates realistic synthetic events from intensity frames and addresses non-idealities such as Gaussian event threshold mismatch and intensity-dependent noise using handcrafted modules. Palinauskas et al. (2023) introduce an event simulator for a robotics use case by extending ESIM (Rebecq et al., 2019) into MuJoCo (Todorov et al., 2012) platform-based simulation. By using different levels of frame interpolation method, Ziegler et al. (2023) introduce an event camera simulator that can perform near-real-time simulation of event camera data.

Physical based simulators.Han et al. (2024) and Lin et al. (2022a) try to incorporate physical law-based imperfections in the event sensors in addition to the changes in the intensity to simulate close-to-real event data streams. By taking into account the fundamental circuit properties, Lin et al. (2022a) develop a realistic event camera simulator that can faithfully model the voltage variations, randomness in photon receptions and noise caused by leakage current into a stochastic process and show that the simulated events strongly resemble real event data. By directly interfacing with a 3D scene, Han et al. (2024) design a realistic lens simulation block and a novel multispectral rendering block to combine both optical as well as circuit-level imperfections. The results show that the system is able to produce high-fidelity data under different lighting conditions and motion speeds.

Learning based simulators.Gu et al. (2021b) and Zhang et al. (2024g) aim to leverage developments in deep learning to approximate the event generation process and hence make it a completely data-driven approach. This removes the elaborate manual efforts required for heuristically tuning each of the components in a traditional simulator, as well as offers more adaptiveness to more domains. Gu et al. (2021b) introduce a method to learn pixel-wise distributions of event contrast thresholds for a given domain, enabling stochastic sampling and parallel rendering to generate event representations that closely match real event camera data. This is achieved through a novel divide-and-conquer discrimination scheme that adaptively assesses synthetic-to-real consistency based on local image and event statistics. This method has also been shown to have better domain adaptation capabilities. V2CE (Zhang et al., 2024g) introduces a two-stage method for realistic event data simulation using a 3D UNet backbone for predicting event voxels, followed by an event sampling module. The whole system is trained using specialized loss functions to enhance the quality of generated event voxels and is shown to be able to convert video to high-fidelity event streams with precise events.

Most of the video restoration and 3D reconstruction models use the event data to fuse it with the RGB sensor data to improve the performance. The characteristics of raw event data are different from conventional RGB camera captures. To achieve better visual media enhancements and reconstruction, the event data needs to be processed for better fusion with image modality. Therefore, researchers developed multiple methods to process the event data better. Another challenge in event cameras is noise and low-resolution capture. Generally, the event data is prone to noise, especially in low-light conditions and therefore, denoising is required. The resolution of event cameras is low as compared to conventional RGB cameras. To fuse the low-resolution event data with high-resolution (HR) RGB camera data, super-resolution event data is necessary. In this section, we will discuss multiple strategies for efficient event data processing, event denoising and event super-resolution in three different sections.

2.4.1 Efficient event feature extraction.

The event camera captures visual information through asynchronous events triggered by brightness changes. The captured event camera data is different from other imaging data. Therefore, special techniques for event cameras are required to process the event data. Huang et al. (2023a) propose a self-supervised method for detecting and describing local features in event streams and introduce a novel event stream representation method called Tencode. Tencode processes event data to obtain pixel-level interest points and descriptors through a neural network. To achieve efficient event data processing, Sun et al. (2022a) introduce the MENet model that features a dual branch structure. Prior methods often ignore the motion continuity between adjacent windows, which leads to the loss of dynamic information and the extraction of more redundant information. MENet addresses these issues by enhancing feature extraction and reducing redundancy, making it a more efficient solution for event stream processing. This dual-branch includes a base branch for full-sized event point wise processing and an incremental branch that captures temporal dynamics between adjacent spatiotemporal windows. The incremental branch, which is equipped with a point wise memory bank, significantly reduces computational complexity and improves processing speed. Traditional event-based backbones often rely on image-based designs, which overlook the unique properties of event data, such as time and polarity. To address this, Peng et al. (2023) introduce group event transformer (GET), a novel vision transformer backbone specifically designed for event-based vision tasks, which decouples temporal-polarity information from spatial information throughout the feature extraction process to use the temporal and polarity information of events fully. It introduces a new event representation named group token, which groups asynchronous events based on their timestamps and polarities. The advantage of this new representation is that it decouples temporal-polarity information from spatial information throughout the feature extraction process. By decoupling temporal and polarity information at the token level, GET enables the Transformer’s attention mechanisms to operate more effectively across both spatial and temporal-polarity domains. This structured representation allows the event dual self-attention block to selectively attend to meaningful patterns. To address the challenge of poor generalizability in existing deep neural networks when deployed at higher inference frequencies of events, Zubic et al. (2024) introduce a novel approach to event-based vision using state-space models (SSMs) with learnable timescale parameters. Transformers typically rely on fixed temporal windows and dense input representations. They struggle with generalization when deployed at different frequencies than those used during training. Transformers are less effective at capturing the continuous temporal dynamics inherent in event streams, leading to performance drops when the input frequency changes. Here, frequency refers specifically to the event representation sampling rate. This is the rate at which the asynchronous event stream is aggregated into a representation for processing. For example, a 20 Hz frequency means events are grouped into 50 ms windows. Higher frequencies (e.g. 100 Hz or 200 Hz) mean smaller time windows (10 ms or 5 ms), which allow for finer temporal resolution but also pose challenges like aliasing. SSMs are designed to handle temporal sequences by maintaining a hidden state that evolves over time. This hidden state captures the underlying dynamics of the event stream, which allows the model to adapt to varying frequencies without needing retraining. The learnable timescale parameters in SSMs help to generalize better across different temporal resolutions and address the poor generalization issue seen in other models. Unlike traditional methods that convert event data into dense image-like representations, FARSE-CNN (Santambrogio et al., 2024) maintains the inherent sparsity and asynchronous nature of event data for efficient asynchronous event processing. The architecture combines recurrent and convolutional neural networks (CNNs) along with compression modules to learn hierarchical features in both space and time. It captures the dynamics of event streams while maintaining sparsity. FARSE-CNN can be considered an event representation technique, as it processes and represents event data in a way that preserves its spatio-temporal sparsity.

2.4.2 Event data training.

Multiple training strategies have been developed to process event information and provide efficient training. Gallego et al. (2019) introduce a collection and taxonomy of 22 objective functions, termed focus loss functions, to analyze events’ alignment with intensity information in motion compensation approaches. The applicability of these loss functions is demonstrated across multiple tasks, including rotational motion, depth and optical flow estimation, showcasing the potential of event cameras in various challenging scenarios. Sparse events are a challenge and they often result in incomplete data, loss of crucial information and difficulties in making accurate predictions. Sparse event completion is essential to enhance data quality by filling in the gaps, leading to more comprehensive and reliable data sets. This improves the accuracy of analysis and predictions, enabling better decision-making. Zhang et al. (2024a) address the challenge of sparse event data. The proposed method treats event streams as 3D event clouds in the spatiotemporal domain and uses a diffusion-based generative model to generate dense event clouds in a coarse-to-fine manner. Yang et al. (2023a) propose a self-supervised learning framework through contrastive learning (Chen et al., 2020; Gao et al., 2021; He et al., 2020; Khosla et al., 2020) to pretrain a network using paired event camera data and natural RGB images for handling event camera data. Contrastive learning is an unsupervised machine learning technique where models learn to differentiate between similar and dissimilar pairs of data by maximizing the similarity of positive pairs and minimizing the similarity of negative pairs to understand and organize data without labeled examples. Yang et al. (2023a) use the pre-trained network in multiple event-driven downstream tasks and show promising performance.

There is a challenge of annotated data scarcity in event-based vision due to the recency of event cameras. To overcome this, Jian and Rostami (2023) propose an unsupervised domain adaptation algorithm to reduce the distributional domain gap between frame-based annotated data sets and event-based data. The domain gap arises because event-based data and frame-based data have different characteristics and distributions. The paper aims to bridge this gap by using unsupervised domain adaptation techniques, allowing knowledge from frame-based annotated data sets to be transferred to event-based data for tasks such as image classification. Their technique tries to map the data from both domains into a shared domain-agnostic embedding space. This algorithm leverages contrastive learning and uncorrelated conditioning to transfer knowledge from conventional camera-annotated data to event-based data. The labeled event data is costly and labor-intensive to annotate. Therefore, Klenk et al. (2024) introduce masked event modeling, a self-supervised learning framework, to reduce the dependency on labeled event data. The proposed method pre-trains a neural network on unlabeled events from any event camera recording. This pretraining significantly improves the accuracy of downstream tasks when the model is fine-tuned. Gehrig and Scaramuzza (2022) explore the necessity of HR sensors in event cameras. The study reveals that, contrary to popular belief, low-resolution event cameras can outperform HR ones in specific conditions, such as low illumination and high-speed scenarios. This is due to the higher per-pixel event rates in HR cameras, which lead to increased temporal noise under these conditions. Lu et al. (2025) introduce the first event-RAW paired and pixel-level aligned data set for event-based image signal processing (ISP). This data set includes 3373 RAW images with a resolution of 2248×3264 and their corresponding events, spanning 24 scenes with three exposure modes and three lenses. The approach involves using the event data to guide the ISP pipeline, which includes steps like demosaicing, white balancing, denoising and color correction. By integrating event data with RAW images, the ISP process can leverage the high temporal resolution and dynamic range of event data to enhance the quality of the resulting RGB images.

Yao et al. (2024) investigate the susceptibility of SNNs to adversarial attacks, specifically targeting raw event data from dynamic vision sensors (DVSs). These sensors capture visual information through asynchronous spikes triggered by brightness changes, offering high temporal resolution and low power consumption. The study introduces a novel adversarial attack approach that directly targets raw event data, addressing the challenges of three-valued optimization and the need to preserve data sparsity. The proposed method treats discrete event values as probabilistic samples, focuses on specific event positions to enhance attack precision and uses a sparsity norm to maintain the original data’s sparsity.

2.4.3 Event spatial and temporal up-conversion.

In low-brightness or slow-moving scenes, events are often sparse and noisy, which poses challenges for event-based tasks. To solve these challenges, Xiang et al. (2022) propose an event temporal up-sampling algorithm that generates more effective and reliable events by estimating the event motion trajectory using a contrast maximization algorithm and then up-sampling the events through temporal point processes.

The spatial low-resolution events directly impact the performance of event-based tasks like reconstruction, detection and recognition. The spatial super-resolution of events leads to enhanced performance in downstream applications such as event-based visual object tracking and object detection. It also helps to improve the quality of the reconstructed images. Li et al. (2021) propose a real-time framework based on a spiking neural network (SNN) to generate HR event streams from low-resolution inputs and deploy it on a mobile platform. SNNs for event data are useful due to their alignment with the asynchronous, sparse nature of event streams. SNNs process data using discrete spikes, making them ideal for capturing fine-grained temporal dynamics. Their energy efficiency and ability to model spatiotemporal patterns enable effective super-resolution of low-resolution event streams. Shariff et al. (2024) integrate binary spikes with Sigma Delta Neural Networks (SDNNs), leveraging a spatiotemporal constraint learning mechanism designed to simultaneously learn the spatial and temporal distributions of the event stream. Zhang et al. (2024d) introduce a self-supervised super-resolution prototype that adapts to any low-resolution input source without requiring prior training or side knowledge. This method leverages asynchronous streaming events to estimate HR counterparts, significantly improving visual richness and clarity. Huang et al. (2024) propose a bilateral event mining and complementary network (BMCNet) that leverages a two-stream network to process positive and negative events individually. This approach allows for comprehensive mining of each event type and facilitates the exchange of information between the two streams through a bilateral information exchange (BIE) module. Zhang et al. (2024d) introduce a unified neural network (CZ-Net) designed to address motion blur using low-resolution events in neuromorphic cameras. The network is a multi-scale blur-event fusion architecture that leverages the scale-variant properties of events and images to achieve cross-enhancement. This architecture effectively fuses cross-modal information, using attention-based adaptive enhancement and cross-interaction prediction modules to mitigate distortions in low-resolution events. Unlike traditional methods for event stream super-resolution, which often require high-quality, HR frames, Weng et al. (2022) propose a recurrent neural network (RNN) based method that does not rely on image frames and achieves super-resolution for a large-scale factor (⁠×16⁠). Traditional methods often mix positive and negative events directly, leading to a loss of detail and inefficiency. Therefore, Liang et al. (2024c) introduce a recursive multi-branch information fusion network (RMFNet) to enhance the spatial resolution of event streams. The method separates positive and negative events to extract complementary information, followed by mutual supplementation and refinement.

2.4.4 Event denoising.

Event cameras, which capture dynamic scenes with high temporal precision, offer advantages like HDR, low latency and low power consumption. However, they are prone to noise due to their differential signal output and logarithmic conversion. They often suffer from noise, particularly under low illumination and varying camera settings. The outputs of event cameras are prone to various types of noises such as photon shot noise, dark current shot noise, leakage current noise and hot pixel noise (Jiang et al., 2024a). Therefore, event denoising becomes crucial for better visual media restoration and 3D reconstruction. In this direction, EventZoom (Duan et al., 2021) addresses the challenge of joint denoising and super-resolving neuromorphic events by leveraging a 3D U-Net backbone architecture, which is trained in a noise-to-noise fashion (Lehtinen et al., 2018) to enforce noise-free event restoration. In the noise-to-noise approach, both input and output are unfiltered noisy data of the same signal and the target is to recover the signal from those noisy measurements. If the noise is zero-mean and independent across samples, then the expected value of a noisy image is the clean image. So, if we train a model to minimize the difference between two noisy versions of the same image, the model learns to predict the underlying clean signal. Duan et al. (2023) also use a 3D U-Net backbone neural architecture to train in a noise-to-noise fashion (Lehtinen et al., 2018). Shi et al. (2024) introduce a novel dual-stage denoising method for event cameras. The proposed method, called polarity-focused denoising (PFD), leverages the consistency of polarity and its changes within local pixel areas to handle noise effectively. Due to camera motion or dynamic scene changes, the polarity and its variations in triggered events are closely linked to these movements, which help to achieve effective noise management. Zhang et al. (2024d) present an iterative coarse-to-fine approach, where an event-regularized prior provides high-frequency structures and dynamic features for blind deblurring, while image gradients help regulate noise removal.

Jiang et al. (2024a) propose the EDformer model that leverages the transformer architecture to learn spatiotemporal correlations among events, enabling effective denoising across varied noise levels. They introduce the ED24 data set, which encompasses 21 noise levels and noise annotations, providing a robust foundation for evaluating denoising algorithms. The data corresponding to 21 different noise levels are defined based on 21 different illumination conditions. Duan et al. (2023) also introduce a display-camera system for recording high frame-rate videos at multiple resolutions for data collection and they trained a 3D-Unet for joint super-resolution and denoising. Display-camera system refers to a setup where a display screen shows high-frame-rate videos and an event camera is positioned to record the display. Duan (2024) also present a comprehensive data set, the LED data set, designed to address the challenges of denoising event camera data in real-world scenarios. They also propose a novel denoising framework, DED that uses homogeneous dual events to better separate noise from signal events. Dual events refer to a dual-sampling setup, where two identical event cameras are used to simultaneously capture the same scene under identical conditions. This setup is designed to exploit the inconsistency of noise events and the consistency of signal events across two recordings. Existing data sets and evaluation metrics for denoising are limited in scale and noise diversity, often relying on sensor information or manual annotation. To address these limitations, Ding et al. (2023) present the E-MLB data set, which includes 100 scenes with 4 noise levels, making it 12 times larger than the largest existing data set. They also introduce a comprehensive benchmark designed to evaluate denoising algorithms for event-based cameras and propose a novel non-reference denoising metric called the event structural ratio (ESR), which measures the structural intensity of events independently of the number of events and projection direction.

2.4.5 Multi-modal and security processing.

Lin et al. (2023a; 2023b) introduce E2PNet, the first learning-based method for registering event data to 3D point clouds. A 3D point cloud is a collection of data points in a three-dimensional coordinate system, representing the external surface of objects or scenes, typically constructed from a 3D scene using techniques like light detection and ranging (LiDAR) or depth sensors. Registering event data to a point cloud is essential for enhancing the accuracy and robustness of scene understanding. Event data, which is more resilient to changes in lighting and motion, when aligned with the spatial information in the point cloud, provides a comprehensive analysis of dynamic scenes. This integration facilitates better object detection, tracking and recognition, making the combined data more reliable and useful for various applications. The core of E2PNet is the event-points-to-tensor network, which encodes event data into a 2D grid-shaped feature tensor. This representation allows the use of established RGB-based frameworks for event-to-point cloud registration without altering hyperparameters or training procedures. Lin et al. (2022b) address the challenge of focus control in event cameras. Traditional autofocus methods are ineffective for event cameras due to differences in sensing modality, noise and temporal resolution. To overcome these challenges, the paper introduces a novel event-based autofocus framework that includes an event-specific focus measure called event rate and a robust search strategy known as event-based golden search. To achieve better restoration and reconstruction, the event data must be aligned to RGB/intensity sensor data. Therefore, Gu et al. (2021c) introduce a novel model for event-camera data alignment. This method is particularly effective for extracting camera rotation, leading to improved event alignment.

EventGPT (Liu et al., 2025c) is the first multimodal large language model for event stream understanding. It integrates the event-driven vision into the large language model. The EventGPT contains an event encoder, a spatio-temporal aggregator, a linear projector, an event-language adapter and an large language model (LLM). To overcome the domain gap between LLMs and event information, they adopt three-stage strategies. At the first stage, GPT-generated RGB image-text pairs warm up the linear projector. The second stage uses the NImageNet-chat data set, a large synthetic data set of event data and corresponding texts, to train the spatio-temporal aggregator and event-language adapter. At the final stage, event-chat, which contains extensive real-world data, is used to fine-tune the entire model and enhance its generalization ability. Each stage reduces the domain gap between events and LLMs to train EventGPT.

To explore vulnerabilities in event data processing systems, Wang et al. (2024b) explore the potential risks of backdoor attacks in event-based vision tasks, which have been under-researched despite the increasing use of asynchronous event data. A backdoor attack is a method of bypassing normal authentication procedures to gain unauthorized access to a system. This type of cyber attack involves exploiting system vulnerabilities or installing malicious software that creates an entry point for the attacker. These event-based vision systems can be compromised by injecting malicious triggers into the event data streams, which can then activate the backdoor during inference, leading to incorrect or manipulated outputs. The authors propose the event Trojan framework, which includes two types of triggers: immutable and mutable. These triggers are based on sequences of simulated event spikes that can be easily incorporated into any event stream to initiate backdoor attacks. The immutable trigger remains constant, while the mutable trigger uses an adaptive learning mechanism to maximize its aggressiveness. To enhance stealthiness, a novel loss function is introduced to minimize the difference between the triggers and original events while maintaining their effectiveness.

High temporal resolution of event camera data enables us to capture the change in scenes at a significantly higher rate than traditional RGB cameras can and this can be used to improve temporal aspects of the RGB images/videos. In this section, we present how to exploit event stream data for video reconstruction and video frame interpolation (VFI). In Figure 3, we provide an overview of how these two methods typically work.

Figure 3.
A flow diagram of event-based video reconstruction from event stream and low frames per second input to high frames per second output.The flow diagram presents an event based video reconstruction pipeline that converts an event stream and low frames per second intensity video into a high frames per second video. The upper path shows event to video conversion where an event stream is transformed into constructed frames. The lower path combines low frames per second video and event stream inputs that pass through flow estimation and synthesis modules. The intermediate outputs are merged in a fusion stage followed by refinement. The final stage produces a high frames per second video with temporally denser reconstructed frames derived from both event and intensity information.

Temporal restoration applications. (Top) Event to video generation models convert a spatio-temporal stream of events with microsecond temporal resolution into a high-quality video, which is typically monochromatic (however, colored outputs also exist through generative coloring). This enables applications such as high-framerate videos and high dynamic range capture. Usually, these models take in a 3D tensor of events along with the feedback from the previous predicted frame to generate the next video frame. (Bottom) Frame interpolation models leverage the high temporal resolution of event streams to estimate non-linear motion information between the frames and insert latent frames between two consecutive frames. As we discuss below, these models can be trained with full supervision, weak supervision or unsupervised methods. The general approach for such models follows processing the RGB and event streams separately and then fusing the information, followed by refinement

Figure 3.
A flow diagram of event-based video reconstruction from event stream and low frames per second input to high frames per second output.The flow diagram presents an event based video reconstruction pipeline that converts an event stream and low frames per second intensity video into a high frames per second video. The upper path shows event to video conversion where an event stream is transformed into constructed frames. The lower path combines low frames per second video and event stream inputs that pass through flow estimation and synthesis modules. The intermediate outputs are merged in a fusion stage followed by refinement. The final stage produces a high frames per second video with temporally denser reconstructed frames derived from both event and intensity information.

Temporal restoration applications. (Top) Event to video generation models convert a spatio-temporal stream of events with microsecond temporal resolution into a high-quality video, which is typically monochromatic (however, colored outputs also exist through generative coloring). This enables applications such as high-framerate videos and high dynamic range capture. Usually, these models take in a 3D tensor of events along with the feedback from the previous predicted frame to generate the next video frame. (Bottom) Frame interpolation models leverage the high temporal resolution of event streams to estimate non-linear motion information between the frames and insert latent frames between two consecutive frames. As we discuss below, these models can be trained with full supervision, weak supervision or unsupervised methods. The general approach for such models follows processing the RGB and event streams separately and then fusing the information, followed by refinement

Close Figure 3.

The event-to-video (E2V) problem is a crucial area of research that focuses on reconstructing traditional intensity images or videos from the asynchronous event streams captured by event cameras (Wang et al., 2019, 2024a). This reconstruction is vital for bridging the gap between event-based vision and conventional frame-based computer vision algorithms. While event cameras offer significant advantages over traditional cameras, the asynchronous structure of output makes directly applying existing computer vision algorithms to event data streams challenging (Rebecq et al., 2019). Early methods emphasized the representational similarities between events and gradients (Cook et al., 2011; Kim et al., 2008; Munda et al., 2018) or optical flow (Bardow et al., 2016). However, these approaches often fell short of achieving realistic reconstructions due to a lack of long-term data and prior information exploration. More recently, deep learning techniques, particularly RNNs and CNNs, have been applied to leverage temporal context for improved performance.

Handcrafted priors and model based approaches. Early attempts focus on visually interpreting or reconstructing intensity images from pure events. Cook et al. (2011) propose using recurrently interconnected areas or maps, to interpret intensity and optic flow. Kim et al. (2008) work with pure events on rotation-only scenes to track the camera and build super-resolution mosaics based on probabilistic filtering. This is later extended to 3D reconstruction and 6-DoF camera motion (Kim et al., 2016). Bardow et al. (2016) reconstruct intensity images and motion fields for generic motion using a variational optimization framework. While applicable to dynamic scenes, these methods often require hand-crafted regularisers, which could lead to a loss of detail in reconstructions. A variational denoising framework is proposed by Munda et al. (2018) and it iteratively filters incoming events based on their timestamps to reconstruct images. However, these direct event integration methods often suffer from “bleeding edges” due to non-uniform contrast thresholds and “ghosting effects” from the unknown initial image. Scheerlinck et al. (2018) propose a computationally efficient asynchronous filter that continuously fuses image frames and events into a single high-temporal-resolution and high dynamic-range image state.

CNN/RNN based approaches. The advent of deep learning significantly advances E2V reconstruction, moving beyond restrictive assumptions. These methods typically group events into spatio-temporal representations like 3D voxel grids or event images to be processed by CNNs (Rebecq et al., 2019; Stoffregen et al., 2020). Mostafavi et al. (2018) used conditional generative adversarial networks to generate HDR images and very high frame rate videos from pure events. E2VID (Rebecq et al., 2019) proposes a novel recurrent network architecture (U-Net-based) to reconstruct videos from a stream of events. They train it on large synthetic data sets and show remarkable generalization to real events, outperforming state-of-the-art reconstruction methods by a large margin in terms of image quality. E2VID+ (Stoffregen et al., 2020) improves performance by better matching synthetic training data statistics to real-world data. SPADE-E2VID (Cadena et al., 2021) introduces spatially adaptive denormalization (SPADE) layers into the E2VID architecture, which improves the quality of early reconstructed frames but increases computational cost. Zou et al. (2021) focus on reconstructing high-speed and high-dynamic range videos and notably collected a paired event and image data set using a coaxial imaging system, though it was not suitable for high-quality nighttime ground truth generation. HyperE2VID (Ercan et al., 2024) proposes a dynamic neural network architecture for event-based video reconstruction that uses hypernetworks to generate per-pixel adaptive filters guided by a context fusion module that combines information from event voxel grids and previously reconstructed intensity images.

Transformer based networks. ET-Net (Weng et al., 2021) is the first to explore the application of transformers for high-speed video reconstruction from an event camera. They propose a hybrid CNN-transformer architecture (ET-Net) that leverages both fine local information from CNN features and global contexts from Transformers, achieving superior reconstruction quality. Furthermore, a token pyramid aggregation strategy to implement multi-scale token integration for relating internal and intersected semantic concepts in the token space is also introduced.

Spiking neural networks. EVSNN/PA-EVSNN (Zhu et al., 2022) is the first to propose image reconstruction using a deep SNNs architecture. They propose the event-based video reconstruction framework based on a fully spiking neural network (EVSNN) and a hybrid potential-assisted framework (PA-EVSNN) using leaky-integrate-and-fire (LIF) and membrane potential (MP) neurons. Their work aims for greater computational efficiency on event-driven hardware. Tang et al. (2025) propose a spike-temporal latent representation (STLR) model for SNN-based E2V reconstruction. This model uses cascaded SNNs, including a spike-based voxel temporal encoder and a U-shape SNN decoder, to solve the temporal latent coding of event voxels for video frame reconstruction, focusing on energy efficiency.

Self-supervised learning (SSL) methods.Paredes-Vallés and de Croon (2020) introduce a self-supervised learning framework that estimates optical flow and reconstructs intensity images via photometric constancy. This approach aims to eliminate the dependency on labeled synthetic data, though it still relies on optical flow estimation, which can be prone to errors. EvINR (Wang et al., 2024a) proposes a novel self-supervised learning approach that re-frames E2V reconstruction as directly solving the event generation model, which is described as a partial differential equation. They use implicit neural representations (INRs) (Chen et al., 2021; Sitzmann et al., 2020) to predict intensity values from spatiotemporal coordinates, eliminating the need for labeled data or optical flow estimation and exhibiting high noise tolerance.

Diffusion models.Liang et al. (2024b) pioneer the use of diffusion models for color video reconstruction from achromatic events. This method addresses the ill-posed nature of E2V (one-to-many mapping) and the regression-to-mean issue by sampling various plausible reconstructions. It uses pretrained diffusion models, a spatiotemporally factorized event encoder and an event-guided sampling mechanism to ensure faithfulness to original events. Zhu et al. (2025) introduce a novel framework that leverages temporal and frequency-based event priors to guide a denoising diffusion probabilistic model (Ho et al., 2020). It uses a temporal domain residual image as the target for the diffusion model and incorporates three conditioning modules: low-frequency intensity estimation, temporal recurrent encoder and attention-based high-frequency prior enhancement. This approach aims to mitigate over-smoothing and blurry artifacts.

Other approaches.Cao et al. (2024) focus on static scene recovery from event cameras by uniquely leveraging illuminance-dependent noise characteristics (photon noise events). It treats noise events as a signal to recover static scene intensity, a component otherwise invisible to event cameras and can complement E2V methods for scenes with both static and dynamic components. EvTemMap (Bao et al., 2024) achieves event-to-dense intensity image conversion using a stationary event camera in static scenes with a transmittance adjustment device (AT-DVS). This method measures the time of event emission for each pixel to form a “Temporal Matrix”, which is then converted to an intensity frame using a neural network, aiming to capture diverse scenes beyond dynamic ones. Nighttime event to video reconstruction network (NER-Net) (Liu et al., 2024c) specifically addresses nighttime dynamic imaging, dealing with temporal trailing characteristics and spatial non-stationary distributions of events in low-light conditions. They introduce a learnable event timestamps calibration module (LETC) and a non-uniform illumination aware module (NIAM) for this purpose. EventMamba (Ge et al., 2025) observes that the local event relationship in the spatio-temporal domain is key for video restoration tasks and proposes a method comprising random window offset (RWO) in the spatial domain and a new consistent traversal serialization approach in the spatio-temporal domain, which enhances Mamba architecture’s capability to handle event data stream.

To summarize, E2V reconstruction focuses on creating traditional intensity videos from asynchronous event camera streams, bridging the gap between neuromorphic sensing and conventional vision by leveraging event cameras’ high temporal resolution, dynamic range and lack of motion blur. Approaches have evolved from early handcrafted or model-based methods to advanced deep learning paradigms, encompassing CNNs, RNNs, transformers, SNNs, self-supervised learning and diffusion models, each striving for enhanced reconstruction quality and computational efficiency across various scene conditions.

Video processing tasks such as VFI and motion deblurring (MD) are crucial for enhancing the quality and temporal resolution of digital videos. Traditional frame-based cameras face inherent limitations, including motion blur due to long exposure times and the loss of inter-frame information caused by slow shutter speeds. These limitations often lead to degraded video quality, especially in dynamic scenes or low-light conditions. Event cameras offer a complementary solution to this challenge. Their ability to capture continuous, non-redundant information about local brightness changes makes them highly advantageous for recovering precise motion cues and filling temporal gaps that conventional cameras miss.

3.2.1 Video frame interpolation.

VFI aims to synthesize non-existent intermediate frames between consecutive frames, thereby increasing the video’s frame rate. Many VFI methods traditionally rely on linear motion assumptions, which often fail in complex real-world scenarios. Event cameras offer a powerful alternative by providing dense temporal information to model non-linear motions more accurately.

Flow based approaches. Flow-based methods typically estimate optical flow between consecutive frames and use this flow to warp input frames to generate intermediate ones. Flow-based approaches can be broadly classified into two categories: (i) Flow estimation from only the event stream and (ii) Flow estimation from both the event and the RGB stream:

  1. Flow from only event stream. TimeLens (Tulyakov et al., 2021) proposes a learning-based framework that combines warping-based and synthesis-based interpolation. Its warping module estimates optical flow from event sequences, allowing it to handle motion blur and non-linear motion, unlike traditional methods that compute flow from frames and assume linear motion. TimeLens++ (Tulyakov et al., 2022) is an extension of TimeLens, improving efficiency and performance by encoding optical flow as cubic splines and using multi-scale feature fusion (MSFF). It computes a parametric motion model from boundary frames and inter-frame events. TimeReplayer (He et al., 2022) introduces an unsupervised learning framework for event-based video interpolation. It directly estimates optical flows between an intermediate frame and input frames using event streams, which helps break the uniform linear motion assumption commonly found in frame-only methods. A2OFWu et al. (2022) propose an end-to-end training method that uses events to generate optical flow distribution masks. These masks provide anisotropic weights for blending and generating intermediate optical flow in orthogonal directions, allowing for a better description of complex motion than isotropic methods.

  2. Flow estimation from both the event as well as the RGB stream. EIF-BiOFNet (Kim et al., 2023) emphasizes direct estimation of asymmetric inter-frame motion fields by effectively leveraging the distinct characteristics of both events and images, avoiding reliance on approximation methods. It also incorporates an interactive attention-based frame synthesis network. IDO-VFI (Shi et al., 2023) estimates optical flow using both frames and events. It then uses a Gumbel gating module to dynamically decide whether to compute additional residual optical flow in specific sub-regions based on the optical flow amplitude, aiming to reduce computational overhead while maintaining high quality. Chen et al. (2023b) suggest that event sequences are better used to calibrate optical flow estimated from RGB images rather than for direct optical flow estimation from events alone. They propose an event-guided recurrent warping strategy and a proxy-guided synthesis strategy to better exploit the quasi-continuous nature of event signals. Ma et al. (2024) further advance event-based VFI by estimating nonlinear per-pixel motion trajectories between two input frames, enabling it to model more complex motions than previous approaches. It achieves this through iterative motion estimation. TimeTracker (Liu et al., 2025b) proposes a continuous point tracking method consisting of a scene aware region segmentation module and a continuous trajectory guided motion estimation module, which is finally used in global motion optimization followed by frame refinement. The experiments show the superiority of TimeTracker in fast nonlinear motion scenarios.

Synthesis based approaches. Synthesis-based methods directly regress intermediate frames or fuse information without explicitly relying on optical flow estimation or image warping operations. Liu et al. (2024b) propose a purely synthesis-based E-VFI framework. It directly synthesizes interpolated frames by aligning key frames to an event-based reference that encodes structural information, thereby circumventing the challenges of non-linear motion fitting and occlusion issues typically introduced by warping operations. Gao et al. (2022a) introduce a fast-slow joint synthesis framework for high-speed VFI. It divides the task into two sub-tasks, focusing on content with and without high-speed motions, tackled by a fast synthesis pathway (including an SNN-based hybrid module for high-speed content) and a slow synthesis pathway, respectively. RE-VDM (Chen et al., 2025) proposes to adapt pretrained video diffusion models for frame interpolation to overcome the data scarcity problem. By using per-tile denoising and two-sided fusion, it achieves better interpolation consistency and generalization on unseen real-world data. To tackle the problem of latency in interpolating between two frames, Musunuri et al. (2024) propose to extrapolate RGB frames from the initial RGB frame and asynchronous events.

Hybrid and attention based approaches. Hybrid approaches combine elements of both flow-based and synthesis-based approaches or heavily rely on attention mechanisms and feature fusion for robust interpolation. Kilicc et al. (2022) propose a lightweight, kernel-based method that fuses event information with standard video frames using deformable convolutions and a multi-head self-attention mechanism. This approach aims to generate crispier frames less vulnerable to blurring and ghosting artifacts. Akman et al. (2023) introduce motion awareness through event-based motion masks, which guide deformable convolutions during image generation to focus on moving areas. It also incorporates a motion-aware loss function to further improve performance. Zhang et al. (2023e) propose a unified operation using inter-frame attention to explicitly extract both motion and appearance information simultaneously. It reuses the attention map for both appearance feature enhancement and motion information extraction, aiming for efficiency. Jin et al. (2022) follow a flow-guided synthesis framework, using optical flow estimation, warping of input frames and context features and synthesizing intermediate frames from these warped representations. Nottebaum et al. (2022) propose a method that fine-tunes an initial block-based principal component analysis basis end-to-end for VFI, creating a general projection space for all images and resolutions and optimizing the representation for the VFI task. Wang et al. (2023b) propose a novel task of event-based continuous color video decompression. It uses a joint synthesis and motion estimation pipeline, where the synthesis module uses a K-plane-based factorization to encode event-based spatiotemporal features and the motion estimation module estimates time-continuous nonlinear trajectories. Cho et al. (2024) introduce a test-time adaptation framework for event-based VFI to address the significant performance drop experienced by event-based VFI methods when applied to different domains (i.e. test data sets with varying distributions from training data). They achieve this by using confident pixels as pseudo ground-truths while mitigating overfitting to recurring scenes by blending historical samples with current inputs. This is crucial for real-world applications where device and environment variations are unavoidable.

3.2.2 Motion deblurring.

MD is a challenging and ill-posed problem due to the loss of motion information and intensity textures during the blur degradation process. Event cameras offer a unique advantage by inherently providing precise motion information and sharp edges, which can be leveraged to alleviate motion blur.

Model based approaches. Model-based approaches derive deblurring solutions based on the physical principles of event generation and image formation, sometimes coupled with optimization techniques. Pan et al. (2019b) propose the event-based double integral (EDI) model to connect intensity images with event data by integrating over events as well as image sequences while considering the blur process. It aims to recover a sharp image and then reconstruct a high frame-rate video, implying a sequential or integrated approach to deblurring and high frame-rate reconstruction. Pan et al. (2019a) extend EDI to multiple event based double integral (mEDI) to analytically reconstruct high frame-rate sharp videos from blurry frames and associated events. This work specifically highlights the challenge of addressing blur in combined event-intensity solutions. E-CIR (Song et al., 2022) extends the EDI model by using a deep learning model to predict a sharp video represented by parametric polynomials from a blurry frame and associated events. It also formulates a refinement objective to encourage temporal propagation of sharp visual features and address temporal smoothness issues.

Learning based approaches. Learning-based methods use neural networks (CNNs, RNNs, attention mechanisms) to learn the mapping from blurry inputs and event data to sharp outputs, often suppressing noise and handling complex real-world conditions. Lin et al. (2020) use a network to estimate intensity residuals between latent sharp and blurry images directly from event integrals, effectively using dynamic filters to handle spatially and temporally variant triggering thresholds. Sun et al. (2022b) propose an event-image cross-modal attention fusion module (EFNet) that jointly extracts and fuses information from event streams and images. They demonstrate that the principled architecture leverages event information more effectively for deblurring compared to simple concatenation or optical flow estimation. Kim et al. (2021) propose a novel exposure time-based event selection module to selectively use event features by estimating the cross-modal correlation between the features from blurred frames and the events and propose a feature fusion module to fuse the selected features from events and blurred frames effectively. Sun et al. (2024) introduce deviation accumulation (DA) as an event preprocessing method to enhance motion perception, enabling the model to distinguish different motion patterns. DA considers the average deviation of polarities at each pixel position over the exposure period. The model also features a recurrent motion extraction (RME) module for multi-scale motion extraction and a feature alignment and fusion (FAF) module to mitigate inter-modal inconsistencies. Kim et al. (2024a) focus on exploiting long-range temporal dependencies in videos. They propose intra-frame feature enhancement through recurrent cross-modal interactions within the exposure time and inter-frame temporal feature alignment to gather information from surrounding adjacent frames. DiffEvent (Wang et al., 2024e) proposes to formulate event-based image deblurring as an image generation problem using diffusion priors for the image and residual, introducing an alternative diffusion sampling framework. CrossZoom (Zhang et al., 2023f) introduces a novel unified neural network (CZ-Net) to jointly recover sharp latent sequences within the exposure period of a blurry input and the corresponding HR events and presents a multi-scale blur-event fusion architecture that leverages the scale-variant properties and effectively fuses cross-modal information to achieve cross-enhancement. Zhang et al. (2023c) introduce a unified processing structure for joint restoration of blurry images and noisy events. It proposes an event-regularized prior for blind deblurring and uses gradient priors from recovered sharp images to supervise event denoising, aiming for robust performance under severe blur and noise. Spiking-convolutional neural network (SC-Net) (Cao et al., 2023) proposes a hybrid SNN and CNN architecture for various video restoration tasks, including deblurring. It features a spiking neural temporal memory for long-term temporal event correlation and a frame-event spatial aggregation (FESA) module for spatial consistency. Zhang et al. (2023d) propose a scale-aware network with a MSFF module for spatially continuous representation and an exposure-guided event representation for arbitrary target latent images, along with a two-stage self-supervised learning framework to generalize deblurring performance across spatial and temporal domains. NED-Net (Cho et al., 2023) specifically tackles deblurring with non-coaxial event cameras. It proposes an attention-based deformable align module for robust feature-level alignment between image and event data, a local score-based aggregation module and a cross-channel interaction module for texture enhancement. St-EDNet (Lin et al., 2023a; 2023b) proposes a coarse-to-fine framework to effectively use and aggregate information from a single blurry image and corresponding event streams (even if misaligned) from a stereo setup. It simultaneously outputs a sequence of sharp images and a disparity map, leveraging parallax to mitigate artifacts caused by misalignment. EGDeblurring (Xie et al., 2025) uses a diffusion model to generate event guidance, which then can be used to deblur RGB images. This enables the usage of event-based deblurring even in cases with RGB-only capture.

3.2.3 Unified deblurring and frame interpolation.

This class of methods explicitly addresses both deblurring and frame interpolation as a unified task or those whose primary goal of high frame-rate video reconstruction from blurry inputs inherently requires both. Lin et al. (2020) explicitly state their aim to reconstruct a sharp video with an increased frame rate from low frame rate blurry videos and corresponding event streams. The proposed algorithm consists of residual estimation, key frame deblurring and VFI components. Zhang and Yu (2022) present a unified framework for event-based MD and frame interpolation for blurry video enhancement. They use a learnable double integral (LDI) network and a fusion network and crucially, a fully self-supervised learning framework that enables training with real-world blurry videos and events without ground truth images. The LDI network is designed to automatically predict the mapping relation between blurry frames and sharp latent images from the corresponding events. Zhang et al. (2023b) propose a unified neural network framework that can “re-expose” a captured photo by adjusting its “neural shutter”. This allows it to handle multiple shutter-related tasks simultaneously, including image deblurring, VFI and rolling shutter (RS) correction. Cheng et al. (2023) focus on the task of continuous-time video extraction from a single blurry image using events. This inherently involves both deblurring the input image and interpolating a high frame-rate sequence, enabling restoration of sharp latent images at arbitrary timestamps. Lu et al. (2023a) propose a novel self-supervised framework that leverages events to guide RS frame correction and VFI in a unified manner. It aims to recover arbitrary frame rate global shutter frames from two consecutive RS frames. Sun et al. (2023) present a general method for event-based frame interpolation that performs deblurring ad hoc, making it applicable to both sharp and blurry input videos. It uses a bidirectional recurrent network and an event-guided channel-level attention fusion module. Weng et al. (2023) study the challenging problem of blurry frame interpolation under blind exposure (i.e. unknown and dynamically varying exposure time) with event cameras. They propose an exposure estimation strategy guided by event streams and a temporal-exposure control strategy for arbitrary-time interpolation. Yang et al. (2024) develop the first method to explicitly address latency correction in improving event-guided deblurring and interpolation tasks. Event timestamps often deviate from actual intensity changes due to sensor latency, noise and illumination variability, making the temporal discrepancy spatially inconsistent and difficult to model. They introduce an event-based temporal fidelity metric to evaluate the sharpness of reconstructed images from latency-corrected events. Lu et al. (2023b) pioneer the simultaneous exploration of RS correction, deblurring and VFI within a one-stage framework using a unified INR. This approach boasts high efficiency and a lightweight model.

To summarize, this section outlines a robust and evolving landscape for event-based video restoration, moving from initial efforts to recover high frame-rate videos to sophisticated unified frameworks addressing multiple image degradation issues simultaneously. While early works focused on either deblurring or interpolation, a clear trend toward jointly addressing MD and VFI has emerged, often integrated with other challenging tasks like RS correction and continuous intensity recovery. Future advancements will likely continue to refine cross-modal fusion techniques, improve generalization to real-world data and develop more efficient and lightweight architectures to enable widespread practical applications.

The event camera excels at capturing scene changes in HDR scenarios at an exceptionally high frame rate without motion blur. These distinctive features of event camera frames contribute significantly to spatial video enhancement. Various spatial video enhancements can be achieved by fusing event camera frames with conventional RGB camera frames. This section will cover topics such as super-resolution and artifact reduction, HDR enhancement, low-light enhancement, occlusion removal, rain removal and focus control. Literature reviews indicate that such spatial enhancements are feasible through the fusion of event camera information. In Figure 4, we provide an overview of how the spatial restoration methods typically work.

Figure 4.
A flow diagram of intensity and event feature extraction with fusion for enhanced intensity frame reconstruction.The flow diagram illustrates a feature processing pipeline that enhances captured intensity frames using event data. The upper branch processes captured intensity frames affected by low resolution, low dynamic range, defocus blurring, low light, rain streaks, and occlusion through intensity feature extraction. The lower branch processes captured event frames containing rich textures, high dynamic range information, and details under low light and occlusion through event feature extraction. The extracted features from both branches are combined in a feature fusion, processing, and reconstruction module. Event driven loss functions are incorporated during training. The output consists of enhanced intensity frames reconstructed using fused intensity and event features.

The typical pipeline of how spatial (in Section 4) enhancement methods work with event information to restore the degraded images/videos

Figure 4.
A flow diagram of intensity and event feature extraction with fusion for enhanced intensity frame reconstruction.The flow diagram illustrates a feature processing pipeline that enhances captured intensity frames using event data. The upper branch processes captured intensity frames affected by low resolution, low dynamic range, defocus blurring, low light, rain streaks, and occlusion through intensity feature extraction. The lower branch processes captured event frames containing rich textures, high dynamic range information, and details under low light and occlusion through event feature extraction. The extracted features from both branches are combined in a feature fusion, processing, and reconstruction module. Event driven loss functions are incorporated during training. The output consists of enhanced intensity frames reconstructed using fused intensity and event features.

The typical pipeline of how spatial (in Section 4) enhancement methods work with event information to restore the degraded images/videos

Close Figure 4.

Event cameras capture small spatial textural details in the event stream and help achieve video super-resolution (VSR) and artifact reduction. Different literature shows unique approaches to fuse the event information along with low-resolution RGB/ intensity frames to reconstruct high-quality output frames. Those approaches can be classified into three different categories, namely, event information fusion, event information alignment for fusion and event information representation for fusion.

Event information fusion. The algorithms discussed in this paragraph take both low-resolution images and event data as input to develop a fusion technology-based novel architecture for super-resolution. Those approaches assume the alignment between events and intensity images. Wang et al. (2020) propose an event-enhanced sparse learning network (eSL-Net), which takes a low-resolution image and an event pair as input and predicts the HR image. The sparse learning network is based on the idea of compressed sensing and dictionary learning, where both low-resolution and HR images are assumed to have sparse representations in learned dictionaries. eSL-Net is designed by unfolding the iterative shrinkage thresholding algorithm (ISTA) into a deep neural network. This makes the network interpretable, as each layer corresponds to an iteration of ISTA. ISTA is a widely used optimization method for solving sparse coding problems, especially in contexts like compressed sensing and image reconstruction. In a single framework, eSL-Net performs deblurring and denoising together while performing super-resolution. Jing et al. (2021) uncover a pivotal insight in VSR. They discover regions with higher temporal frequency – characterized by smaller pixel displacements between consecutive frames – contribute more effectively to the reconstruction of HR textures. These tiny displacements are crucial because they preserve fine-grained details and subtle variations in the scene, which are often lost in lower-frequency regions with larger motion gaps. This observation leads to a paradigm shift in VSR design and motivates the use of an event camera for VSR, as an event camera captures uniform and tiny pixel displacements between neighboring frames. These cameras produce dense streams of events that inherently encode high-frequency motion information, making them ideal for capturing texture-rich regions. They propose an asynchronous interpolation event asynchronous interpolation module in their VSR framework, which effectively fuses event and image features. EvIntSR-Net (Han et al., 2021) converts event data into multiple latent intensity frames to reduce the domain gap between event streams and intensity frames. As the latent domain extracts high-level information, it is relatively easy to bring the event stream latent closer in the latent domain of intensity frames. Kai et al. (2023) introduce an event-driven bidirectional video super-resolution framework, which has an event-assisted temporal alignment module that leverages events to generate nonlinear motion for aligning adjacent frames. Guo et al. (2023) develop EFSR-Net, which contains coupled response blocks (that fuse data from both RGB and event camera. This enables the recovery of detailed textures in shadows. Noise present in event streams hinders the performance of super-resolution. Separate additional event denoising may lead to over-suppression of events. Therefore, Yu et al. (2023) develop an eSL-Net++ based on a dual sparse learning scheme, where both events and intensity frames are modeled with sparse representations to suppress the noise present in the event stream and handle low resolution with motion blurs together. An event shuffle-and-merge scheme is also proposed to extend eSL-Net++ to generate HR and high framerate video from a single blurry frame without additional training. Liu et al. (2024a) propose an end-to-end framework called EGI-SR, which uses three cross-modality encoders to learn both modality-specific and modality-shared features from stacked events and intensity images. This design helps mitigate the negative impact of modality differences and reduces the feature space gap between events and intensity images. A transformer-based decoder is then used to reconstruct the super-resolved image. Unlike traditional VSR methods that focus on motion learning, EvTexture (Kai et al., 2024) uses the high-frequency details captured by event cameras to improve texture regions. MamEVSR (Xiao and Wang, 2025), a Mamba-based deep neural network (Gu and Dao, 2023; Gu et al., 2021a) for event guided-VSR. Their Mamba-driven network offers a global receptive field with linear computational complexity. By doing so, it overcomes the limitations of the CNN and transformers. MamEVSR’s interleaved Mamba (iMamba) block applies multidirectional selective state space modeling for feature fusion and propagation across bi-directional frames while maintaining linear complexity. MamEVSR’s cross-modality Mamba block exploits spatio-temporal information from both event information and the output of the iMamba block and fuses the cross-modality information. Unlike other approaches, which perform super-resolution for a specific scale factor, Lu et al. (2023c) reconstruct HR frames with an arbitrary scale factor. Their framework consists of the spatial-temporal fusion module to learn 3D features from both events and RGB frames, the temporal filter module to extract explicit motion information from events near the queried timestamp to generate 2D features and the spatial-temporal implicit representation module for arbitrary scale super-resolution.

Event information alignment for fusion. The algorithms discussed in this paragraph deal with the alignment problem between events and intensity images, which is a real-world challenge and is crucial for super-resolution. The existing methods discussed above assume strict calibration between event and RGB cameras, which is often impractical for HR devices like dual-lens smartphones and drones. Here, we will discuss the algorithms that develop a method to align event information with intensity information to perform super-resolution. Asymmetric event-guided video super-resolution network (AsEVSRN) (Xiao et al., 2024a) addresses this by incorporating two specialized designs: the content hallucination module and event-enhanced bidirectional recurrent cells. The content hallucination module dynamically enhances event and RGB information, while the bidirectional recurrent cells align and propagate temporal features using event-enhanced flow. In a similar direction, Xiao et al. (2024b) introduce the event AdapTER (EATER), which includes the event-adapted alignment (EAA) unit and the event-adapted fusion (EAF) unit. The EAA unit aligns multiple frames using event streams in a coarse-to-fine manner, while the EAF unit fuses these frames with event data through a multi-scale design.

Event information representation for fusion. Event representation and event feature extraction also play a crucial role in super-resolution. The methods discussed below present different techniques to represent or extract event information that benefits super-resolution performance. Teng et al. (2022) introduce a novel event representation called neural event stack, which encodes comprehensive motion and temporal information while adhering to physical constraints and suppresses noises in the event stream, making it suitable for image enhancement tasks like super-resolution and deblurring. Event-based blurry super resolution network (Zhang et al., 2024b) includes a multi-scale center-surround event representation to exploit intra-frame motion to extract multiscale texture information inherent in events. For better event feature extraction, Cao et al. (2023) develop a SC-Net, which integrates SNNs with CNNs to process asynchronous event data. The proposed method includes a spiking-convolutional layer that extracts temporal features from event streams. Apart from super-resolution, it performs deblurring and deraining.

Traditional HDR imaging techniques often involve merging multiple low dynamic range (LDR) images taken at different exposures. This can be challenging due to over- or under-exposure issues, leading to ghosting artifacts. On the other hand, event cameras capture HDR scenes as intensity maps. The incorporation of LDR RGB frames and event fusion offers a promising solution for capturing high-quality HDR images in diverse lighting conditions. Literature shows multiple novel HDR imaging pipelines (Han et al., 2023, 2020; Li et al., 2024b; Messikommer et al., 2022; Shaw et al., 2022; Weng et al., 2024a; Yang et al., 2023b) that integrate bracketed LDR image data from standard cameras and HDR events from event cameras to reconstruct HDR images. Shaw et al. (2022) develop an event-to-image feature distillation module that translates event features into the image-feature space with self-supervision and use attention and multi-scale spatial alignment modules for fusion. HDRev-Net (Yang et al., 2023b) uses temporal correlations recurrently to suppress flickering effects in the reconstructed HDR video. Weng et al. (2024a) propose a lightweight multi-scale receptive field block that is used for rapid modality conversion from event streams to frames, while a dual-branch fusion module aligns features and removes ghosting artifacts caused by differences in camera positions and frame rates. An exposure-aware framework (Li et al., 2024b) includes an exposure attention fusion module and incorporates a self-supervised loss based on structural priors to enhance details in saturated areas and reduce noise. Exposure attention fusion is a mechanism designed to adaptively fuse features from standard dynamic range images and event streams based on the exposure level of different regions in the image. It uses an exposure mask to guide this fusion process. Guo et al. (2024c) introduce a diffusion-based fusion module and a real-world fine tuning strategy to enhance the generalization of the alignment module on real-world events. They incorporate image priors from pre-trained diffusion models to address artifacts in high-contrast regions and minimize alignment errors. Color events record asynchronous pixel-wise color changes in an HDR. Cui et al. (2024) incorporate color events into the single-exposure HDR imaging pipeline. An exposure-aware transformer module is designed to propagate informative hints from normally exposed LDR regions and event streams to the missing areas. This module includes an exposure-aware mask, which is a learned attention-like mechanism that helps guide the transformer to focus on reliable regions during HDR reconstruction. It suppresses misleading information from saturated regions and enhances the propagation of trustworthy color hints from well-exposed areas. ERS-HDRI (Li et al., 2024a) is designed to enhance remote sensing images, which often struggle with LDR, leading to incomplete scene information. The proposed framework addresses this by using a coarse-to-fine strategy, integrating the event-based dynamic range enhancement (E-DRE) network and the gradient-enhanced HDR reconstruction (G-HDRR) network. The E-DRE network extracts dynamic range features from LDR frames and event streams, performing intra- and cross-attention operations to fuse multi-modal data. A denoise network and a dense feature fusion network generate a coarse HDR image, which is then refined by the G-HDRR network using a gradient enhancement module and a multiscale fusion module. Self-EHDRI (Xiaopeng et al., 2024), a self-supervised learning paradigm, generalizes HDR enhancement performance in real-world dynamic scenarios. This framework uses a self-supervised learning strategy to learn cross-domain conversions from blurry LDR images to sharp LDR images, enabling the recovery of sharp HDR images even without ground-truth sharp HDR images. Unlike existing, AsynHDR (Wu et al., 2024) does not combine RGB frame-based sensors with event sensors in the same system. AsynHDR is a pixel-asynchronous HDR imaging system, a novel capture system that integrates both event-based sensors (DVS) and optical modulation components liquid crystal display (LCD panels). It integrates event sensors with LCD panels that modulate the irradiance incident upon the event sensors, triggering pixel-independent event streams. It achieves HDR imaging solely using DVS, modulated light and a novel temporal-weighted reconstruction algorithm. The system encodes scene brightness into the timing of events, which enables the reconstruction of HDR images with HDR and reduced noise. The authors mention that if a DVS sensor with a Bayer matrix is used, the system could be extended to color HDR imaging, similar to RGB cameras.

Traditional RGB cameras often struggle with long exposure times in low-light conditions, leading to motion blur and reduced visibility. With their HDR and temporal resolution, event cameras can effectively capture motion information even in very dark settings. This unique capability of event cameras helps to achieve low-light enhancement. Zhang et al. (2020) propose a novel unsupervised domain adaptation network that translates HDR events in low light into sharp, canonical images as if captured in daylight. Event cameras naturally provide HDR data, which preserves coarse scene structure even in darkness. The domain adaptation framework leverages this HDR property to bridge the gap between noisy, sparse, low-light events and rich daylight images without any need for paired data. Liu et al. (2024c) develop a NER-Net that includes a LETC to align temporal trailing events and a NIAM to stabilize the spatiotemporal distribution of events. The above-mentioned works mainly convert events in low-light conditions into intensity videos to improve the visibility in low-light scenarios. Other works mainly rely on fusion strategies (Jiang et al., 2023; Liu et al., 2023; Wang et al., 2024d; Yunfan et al., 2025). Incorporation of event information with low-light-intensity videos is used to elevate the performance of the low-light scene enhancement framework. To ensure temporal stability and restore details, Wang et al. (2024d) design unsupervised temporal consistency loss and detail contrast loss, which, along with supervised loss, contribute to the semi-supervised training of the network on unpaired real data. Jiang et al. (2023) introduce a residual fusion module to minimize the domain gap between event streams and frames by using the residuals of both modalities. Here, residuals refer to residual connections or residual blocks commonly used in deep learning architectures, especially in CNNs and transformer-based models. Yunfan et al. (2025) propose SEE-Net, a novel framework for enhancing images captured under a wide range of lighting conditions using event camera data. It takes as input a standard RGB image and its corresponding event stream and outputs a brightness-adjustable image guided by a user-defined brightness prompt. The method begins by embedding spatial and sensor-specific information into both image and event data. These features are fused using cross-attention mechanisms to form a broader light range representation, which captures illumination dynamics across lighting extremes. An MLP decoder then integrates the brightness prompt to generate the final enhanced image, allowing pixel-level control over exposure. This design enables flexible brightness adjustment during inference and robust training across diverse lighting scenarios. They also propose SEE-600K, a large-scale data set comprising 610, 126 image-event pairs collected from 202 distinct scenes, each captured under approximately four different lighting conditions, spanning a dynamic illumination range of over 1,000-fold. Liang et al. (2024a) address the limitations of existing research by introducing a large-scale data set comprising over 30,000 pairs of images and events captured under varying illumination conditions. This data set was meticulously curated using a robotic arm to ensure precise spatial and temporal alignment. The proposed method, EvLight, integrates structural and textural information from both images and events through a multi-scale holistic fusion branch. To handle variations in regional illumination and noise, the authors introduce SNR-guided regional feature selection, enhancing features from high signal to noise ratio (SNR) regions and augmenting those from low SNR regions by extracting structural information from events. The extension of the approach, EvLight++, is developed by Chen et al. (2024a). It extends the method to videos and proposes some modifications, like a ConvGRU recurrent module to capture long-range temporal dependencies, a temporal loss to ensure illumination consistency across frames and new spatio-temporal alignment using a matching strategy. Tian et al. (2025) propose a framework that combines long- and short-exposure frames with event camera data to enhance low-light images. It uses long-exposure illumination to improve short-exposure frames, applies SNR-guided fusion for noise reduction and detail preservation and leverages event data for better feature selection in noisy regions.

Along with low-light enhancement, a few works propose algorithms to handle multiple degradations during low-light scenes. Liang et al. (2023) effectively address challenges like motion blur and noise in low-light conditions. They use hybrid inputs of events and frames to capture temporal correspondences and provide alternative observations, such as intensity ratios between consecutive frames and exposure-invariant information. A neural network is trained to establish spatiotemporal coherence between visual signals of different modalities and resolutions by constructing a correlation volume across space and time. Kim et al. (2024b) develop an end-to-end framework that leverages temporal information from both events and frames, incorporating a cross-modal feature module to enhance structural information while suppressing noise to perform low-light video enhancement and deblurring. They also develop the real-world event-guided low-light enhancement and deblurring data set, which includes synchronized low-light blurred images, normal-light sharp images and low-light event streams. Zhang et al. (2024f) address the challenges of VFI in low-light conditions using event cameras and propose a novel per-scene optimization strategy that leverages the internal statistics of a sequence to handle degraded event data. This approach improves the generalizability of VFI to different lighting and camera settings.

Traditional synthetic aperture imaging (SAI) methods often require prior information and have strict camera motion constraints. They also struggle with dense occlusions and extreme lighting conditions, leading to performance degradation. Event camera helps to overcome these limitations and provides a robust solution for capturing detailed images with significant occlusions, enhancing the capabilities of imaging systems in various applications. REDIR (Guo et al., 2024a), an end-to-end refocus-free variable event-based SAI method. It aligns global and local features of variable event data to achieve occlusion-free imaging from pure event streams without needing prior information (like camera motion parameters or manual focusing). REDIR incorporates a perceptual mask-gated connection module (PMCM) to interlink information between modules and a temporal-spatial attention (TSA) mechanism within the SNN block to enhance target extraction. PMCM connects the registration and reconstruction parts of the network. It filters occlusion events and helps transfer useful features between modules. The TSA mechanism helps the model focus on persistent target signals across time and space, filtering out transient occlusion noise. The event registration module uses a neural network to align event data from different times. It handles camera shake, rotation and zoom automatically. Occlusion filtering filters out noise from occluding objects using a smart attention mechanism. It keeps only the useful signals from the target. Unlike REDIR, other approaches adopt fusion strategies for occlusion removal. Most of the approaches consist of a hybrid network combining SNNs and CNNs. The SNN layers encode the spatio-temporal information from the event data, while the CNN decoder transforms this information into visual images of the occluded targets. CNNs and SNNs-based fusion systems combine the strengths of event cameras and frame-based cameras to generate occlusion-free visual images in both sparse and dense occlusions (Guo et al., 2024b; Li et al., 2022; Liao et al., 2022; Yu et al., 2022; Zhang et al., 2021). The EF-SAI system of Liao et al. (2022) processes multi-modal features through a multi-stage fusion network that enhances cross-modal information and selects density-aware features. Li et al. (2022) present a comparison loss function that is introduced to enhance the clarity of the reconstructed images. They introduce an SNN-based encoder that denoises and encodes the asynchronous event data. The encoded event features and frame features are fused using a custom-designed fusion layer and a joint decoder reconstructs the final clear image. Zhang et al. (2023a) capture continuous streams of events from occluded scenes and integrate event information from continuous viewpoints using a novel cross-view mutual attention mechanism for effective fusion and refinement. For occlusion removal, the event camera is rapidly moved across a scene. This movement allows the camera to capture the scene from slightly different perspectives over time and it effectively simulates multiple views of the same scene. In another work, Zhang et al. (2024b) process these events by integrating multi-view spatial-temporal information through a long-short window feature extractor (LSW) and a cross-view mutual attention-based module for improved fusion and refinement. The LSW is designed to handle the spatial-temporal richness of spike data. Dense window representation preserves fine-grained temporal details. Long window representation accumulates spikes over a longer time to simulate a blurred image-like view, which is helpful to capture structural information. These representations are then fused using a cross-view attention (CVA) module. Guo et al. (2024b) present an event stream encoder based on SNNs, which efficiently encodes and denoises the event data for de-occlusion. Instead of binary spikes, they used full-precision LIF (FP-LIF) neurons, which retain the actual MP values when firing. This prevents information loss during encoding, which is a challenge in traditional LIF. FP-LIF helps to improve feature quality. They also introduce isomorphic network knowledge distillation to solve the limited data issue. Yu et al. (2022) use a refocus-net module, which refocuses collected events to align in-focus events while scattering out off-focus ones. The concept of refocusing is used to address the problem of occlusion. The proposed method refocuses event streams to align signal events from occluded targets while scattering noise events from foreground occlusions. This helps to achieve high-quality reconstruction even under very dense occlusions. Previous methods could only perform at a single depth plane, which limits their usefulness in scenes with multiple depth layers. Liu et al. (2025a) introduce a method to refocus each event individually based on its depth and it enables occlusion-free imaging across multiple depths. They prove that it is feasible to estimate depth maps from multi-view event data even under dense occlusions, which was previously considered very challenging. They introduce a depth estimation module that uses a time-guided attention module to capture temporal and viewpoint relationships. Deformable residual blocks to adaptively extract features across scales. The entire pipeline is trained end-to-end using only occlusion-free multi-view images as supervision. There is no need for ground-truth depth maps.

Video deraining models using events contribute to the field by offering a powerful tool for improving video quality in adverse weather conditions. This network leverages the unique capabilities of event cameras to separate rain streaks from the background, enhancing their utility in challenging weather conditions. Experimental results show that these methods outperform existing deraining techniques for traditional camera videos. These approaches represent a significant advancement in the field of video deraining, offering promising potential for future research and applications. The approach developed by Cheng et al. (2022) operates in the width and time (W–T) space, leveraging the discontinuity of rain streaks in these dimensions while maintaining the smoothness of background objects. “width” refers to the horizontal spatial dimension of the event frame, which is the number of pixels across each row of the image. 3D event video data (height × width × time) is transformed into a W–T space by slicing along the height axis. In this W–T space, each image is a 2D slice where the horizontal axis is the width (i.e. pixel columns) and the vertical axis is time (i.e. event frame index). This transformation helps reveal the discontinuity of rain streaks along the width and time dimensions, as the rain typically moves vertically and affects only a few adjacent pixels in width, it appears as sparse, noise-like patterns in W–T space. This allows raindrops and streaks to be treated as uniform noise, which can be effectively removed using a non-local means filter.

Later, Zhang et al. (2023c)’s end-to-end learning-based network that includes an event-aware motion detection module, which adaptively aggregates multi-frame motion contexts using event-aware masks to capture motion information from event streams effectively and a pyramidal adaptive selection module, which reliably separates background and rain layers by incorporating multi-modal contextualized priors, ensuring accurate rain removal and preservation of background details. As rain steaks can be identified easily in event data, Wang et al. (2023a) leverage this and introduce an end-to-end unsupervised learning-based network for the first time. It consists of two key modules: the asymmetric separation module to segregate features of the rain and background layers and the cross-modal fusion module to enhance positive features and suppress negative ones from a cross-modal perspective. To address the difficulty of modeling the temporal and spatial correlations of rain streaks in videos with existing methods, Sun et al. (2023) propose approach include multi-patch progressive learning, which divides video frames into multiple patches and processes them progressively to capture intricate rain streak details and an event-aware mechanism that leverages event-based sensors to detect and remove rain streaks effectively. Similar to other domains, spiking CNNs have been adopted in deraining to adapt sparse event sequences to capture the features of falling rain (Fu et al., 2024; Ruan et al., 2024). Fu et al. (2024) introduce a bimodal feature fusion module, which combines dense convolutional features from video frames with sparse spiking features from event sequences to enhance the network’s ability to identify and remove rain streaks accurately. Similarly, Ruan et al. (2024) use a spiking network to reconstruct a rain-free background and extract the physical characteristics of rain. Ge et al. (2024) use event signals as prior knowledge to improve dynamic information perception and design a deep unfolding optimization algorithm to construct a de-raining network. The network combines CNNs for spatial feature extraction and SNNs for temporal dynamics. This hybrid design allows the network to leverage both dense intensity frames and sparse event streams. It helps to improve rain removal while preserving textures. Spiking mutual enhancement (SME) module in the network selectively focuses on relevant spatio-temporal regions and enhances dynamic feature extraction from event streams.

The emergence of event cameras, which capture changes in the scene at a high temporal resolution, opens up new possibilities for addressing the challenges of fast and accurate auto-focus in adverse conditions. Bao et al. (2023) propose an event-based focusing algorithm by leveraging the symmetrical relationship between event polarities. The results demonstrate that precise focus, with less than one depth of focus, is achieved within 0.004 s on a self-built high-speed focusing platform. Bao et al. (2025) propose a one-step event-driven autofocus algorithm and they reduce the focusing time and focus error significantly as compared to their prior work. They introduce the event Laplacian product (ELP) focus detection function. It combines event data with grayscale Laplacian information and formulates the autofocus problem as a detection task. The Laplacian of the grayscale image captures image sharpness and edge information. Specifically, it highlights regions of rapid intensity change, which correspond to edges or fine details in the image. In the context of autofocus, sharper images have stronger Laplacian responses, while blurred images have weaker or smoother Laplacian values. ELP is essentially a multiplication of the Laplacian of the grayscale image and the event data, followed by a summation. The sign of the ELP value changes (a “sign mutation”) when the system reaches the focus position. This mutation is used to detect the focus point in real time and it helps to achieve one-step autofocus without scanning through a focus stack. The Laplacian provides spatial texture cues that, when combined with temporal event data, allow the system to determine both the focus position and direction of adjustment.

An event camera also helps in generating all-in-focus images using event signal streams during a continuous focal sweep. Lou et al. (2023) introduce a method to generate high-quality all-in-focus images from a single shot, addressing the challenges of traditional focal stack methods that require multiple shots. They propose the concept of an event focal stack, which consists of event streams captured during a continuous focal sweep. The process involves three main steps: first, automatically selecting the optimal timestamps for refocusing based on the high temporal resolution of event streams; second, using these timestamps and corresponding events to reconstruct a series of refocused images, creating an image focal stack; and finally, merging these refocused images with weights predicted from the images and neighboring events to produce a sharp, all-in-focus image. The extension of this approach is proposed by Teng et al. (2024), which includes multiple improvements like reformulating the all-in-focus imaging pipeline and handling an arbitrary number of image focal stacks for merging.

To summarize, the spatial enhancement domain demonstrates the transformative potential of event cameras in improving visual quality across diverse challenging scenarios, with their ability to capture fine spatial textural details and HDR information making them invaluable complements to traditional RGB sensors. The field has achieved significant breakthroughs across multiple applications: super-resolution and artifact reduction through sophisticated fusion approaches that evolved from basic concatenation to advanced alignment techniques and representation learning; HDR enhancement via integration of event streams with bracketed LDR images, effectively addressing ghosting artifacts and exposure issues through exposure-aware fusion mechanisms and diffusion-based approaches; low-light enhancement leveraging event cameras’ superior dynamic range capabilities for both unsupervised domain adaptation and multi-degradation handling; and specialized applications including occlusion removal through SAI, rain removal by exploiting temporal discontinuity patterns and focus control enabling rapid autofocus systems. The technical evolution has progressed from simple fusion methods to sophisticated architectures incorporating attention mechanisms, SNNs, transformers, diffusion models and state-space models, resulting in more robust and generalizable solutions. These spatial enhancement advances have proven particularly transformative for autonomous systems, computational photography and surveillance applications where traditional cameras fail under extreme conditions such as high-speed motion, challenging lighting, occlusions and adverse weather, establishing event cameras as essential components for next-generation visual media restoration systems.

Event cameras register asynchronous per-pixel brightness changes at high rates, which makes them ideal for observing fast motion with minimal motion blur. Due to the high temporal resolution of the asynchronous visual information acquisition, the output of these sensors is ideally suited for dynamic 3D reconstruction like the N-ocular 3D reconstruction algorithm for event-based vision data (Carneiro et al., 2013), 3D reconstruction using data from a stereo event-camera rig in static scenes (Zhou et al., 2018), reliable 3D hand mesh reconstruction (Jiang et al., 2024b), 3D hand sequence recovery (Park et al., 2024). An event camera has the unique ability to respond to edges of a captured scene, providing geometric information without preprocessing and continuously measures as the sensor moves. Rebecq et al. (2018) exploit this and introduce an event-based multi-view stereo algorithm to estimate semi-dense 3D structures from a known camera trajectory. Baudron et al. (2020) present an event-to-silhouette neural network module that converts event frames into silhouettes and includes neural branches for camera pose regression. Traditional photometric stereo techniques require capturing multiple HDR images under different lighting conditions, which limits their speed and real-time applicability. EventPS (Yu et al., 2024a) overcomes these limitations by estimating surface normal directly from radiance changes detected by the event camera.

Along with the advancement of event cameras’ suitability in different 3D reconstruction domains, photorealistic 3D reconstruction with events gained momentum. Figure 5 shows the typical pipeline of how event information is used to improve 3D reconstruction. Figure 6 shows the overall progress of photorealistic 3D reconstruction of events and the novel research opportunities that still persist. Accurate photorealistic 3D reconstruction often requires the fusion of information from multiple viewpoints. Real-world deployment of neural radiance fields (NeRF) and 3D Gaussian splatting (3DGS) methods faces a set of challenges that impact the fidelity and robustness of scene recovery. These include issues such as synchronization and calibration in multi-camera setups, robustness to noise and lighting variations and limitations in resolution and focus. Event cameras, with their unique sensing modality, provide valuable capabilities to address many of these challenges. As event cameras operate asynchronously and produce per-pixel timestamped events at microsecond resolution, they offer fine-grained temporal information that can be used to achieve highly accurate multi-camera synchronization. This enables consistent spatial alignment and helps prevent artifacts caused by temporal misalignment. Their unique characteristics enable them to complement conventional frame-based methods, particularly in scenarios with fast motion, varying illumination and sensor noise. Recent works have explored the integration of event data into neural representations for photorealistic 3D reconstruction, including NeRF and 3DGS.

Figure 5.
A flow diagram of event guided 3 D reconstruction with intensity and event loss supervision.The flow diagram presents a 3 D reconstruction training pipeline that integrates captured intensity frames and captured event frames. Captured intensity frames pass through an intensity correction stage and are processed by 3 D reconstruction algorithms such as N e R F and 3 D G S to generate a synthesized 3 D scene. The synthesized scene is rendered to produce intensity outputs and a simulated event stream. The simulated event stream contributes to event loss, while rendered intensity contributes to intensity loss. Captured event frames are also converted from event to intensity representations to support intensity loss computation. The pipeline combines rendering, simulation, and dual loss supervision to guide 3 D reconstruction learning.

Typical pipeline for 3D reconstruction using event cameras. The process begins by leveraging the event data to enhance or restore degraded intensity frames. These restored frames provide a more reliable foundation for subsequent processing. Once the intensity information is refined, a 3D reconstruction algorithm is applied to generate a spatial representation of the scene. During the reconstruction phase, novel views of the scene are rendered to simulate different perspectives. The system uses both intensity loss and event loss as guiding signals to optimize the reconstruction quality, ensuring that the output is consistent with both the original intensity frames and the event stream

Figure 5.
A flow diagram of event guided 3 D reconstruction with intensity and event loss supervision.The flow diagram presents a 3 D reconstruction training pipeline that integrates captured intensity frames and captured event frames. Captured intensity frames pass through an intensity correction stage and are processed by 3 D reconstruction algorithms such as N e R F and 3 D G S to generate a synthesized 3 D scene. The synthesized scene is rendered to produce intensity outputs and a simulated event stream. The simulated event stream contributes to event loss, while rendered intensity contributes to intensity loss. Captured event frames are also converted from event to intensity representations to support intensity loss computation. The pipeline combines rendering, simulation, and dual loss supervision to guide 3 D reconstruction learning.

Typical pipeline for 3D reconstruction using event cameras. The process begins by leveraging the event data to enhance or restore degraded intensity frames. These restored frames provide a more reliable foundation for subsequent processing. Once the intensity information is refined, a 3D reconstruction algorithm is applied to generate a spatial representation of the scene. During the reconstruction phase, novel views of the scene are rendered to simulate different perspectives. The system uses both intensity loss and event loss as guiding signals to optimize the reconstruction quality, ensuring that the output is consistent with both the original intensity frames and the event stream

Close Figure 5.
Figure 6.
A conceptual diagram of event driven 3 D reconstruction methods and related components.The diagram illustrates a central module labeled 3 D reconstruction methods such as N e R F and 3 D G S connected to multiple surrounding components. These include single blurry image and events to photorealistic 3 D scene, single moving event camera information to photorealistic 3 D scene, event driven intensity change loss for blur and low light handling, event driven priors such as motion, geometry, and density, and joint optimization of camera pose and 3 D scene reconstruction. Additional components include handling degradations such as super resolution, high dynamic range, and out of focus blur, sparse intensity camera setup, and multi camera synchronization and calibration.

The summary of recent works in photorealistic 3D reconstruction using NeRF and 3DGS with the help of events. The areas marked as green show the already explored areas in NeRF and 3DGS. The red-marked areas show the future potential use cases of NeRF and 3DGS

Figure 6.
A conceptual diagram of event driven 3 D reconstruction methods and related components.The diagram illustrates a central module labeled 3 D reconstruction methods such as N e R F and 3 D G S connected to multiple surrounding components. These include single blurry image and events to photorealistic 3 D scene, single moving event camera information to photorealistic 3 D scene, event driven intensity change loss for blur and low light handling, event driven priors such as motion, geometry, and density, and joint optimization of camera pose and 3 D scene reconstruction. Additional components include handling degradations such as super resolution, high dynamic range, and out of focus blur, sparse intensity camera setup, and multi camera synchronization and calibration.

The summary of recent works in photorealistic 3D reconstruction using NeRF and 3DGS with the help of events. The areas marked as green show the already explored areas in NeRF and 3DGS. The red-marked areas show the future potential use cases of NeRF and 3DGS

Close Figure 6.

Advancement of NeRF models with events.Rudnev et al. (2023) developed the first approach in the domain of dense and photorealistic 3D reconstruction using a single color event camera. They used only event camera data for photorealistic 3D reconstruction. After that, the advancement in the direction of RGB and event camera information fusion has gained momentum to improve the 3D reconstruction performance of the NeRF model using additional event information. This helps to address multiple challenges of 3D reconstruction, achieving both image deblurring and high-quality, sharp NeRF reconstruction from blurry input images, which is common in real-world scenarios due to moving cameras and objects (Klenk et al., 2023; Qi et al., 2023; Rudnev et al., 2024). Qi et al. (2023) introduce a blur rendering loss and an event rendering loss to guide the network by modeling the real blur process and event generation process, respectively. Low and Lee (2024) propose the Deblur e-NeRF method, which addresses this by incorporating a physically accurate pixel bandwidth model to account for event motion blur and introducing a threshold-normalized total variation loss to improve the regularization of large textureless patches. Physically accurate pixel bandwidth refers to a detailed model of how an event camera pixel responds to changes in light intensity over time, based on the actual analog circuitry of the sensor. In low-light or high-speed scenarios, limited bandwidth causes motion blur because the pixels cannot react quickly enough and it leads to delayed or missing events. The Deblur e-NeRF paper models this behavior using a cascade of low-pass filters that simulate realistic event simulation and improved NeRF reconstruction. Li et al. (2024c) explore the possibility of recovering NeRF from a single blurry image and its corresponding event stream. To capture the blurry image, both the RGB camera and the event camera are assumed to follow the same continuous camera motion trajectory during the exposure period. The proposed method jointly learns the NeRF and the camera’s continuous motion trajectory by minimizing the difference between synthesized and real measurements of both RGB camera data and event camera data. Qi et al. (2024b) propose EBAD-NeRF, a method designed to enhance NeRF by addressing motion blur issues in low-light and high-speed scenarios and introduce an intensity-change-metric event loss and a photometric blur loss to model camera motion blur explicitly.

Other than handling the blurry and low-light scenes, the pose estimation is a challenge in NeRF models. Ma et al. (2023) address challenges such as unknown absolute radiance at event locations and unknown camera poses during events. Their proposed method jointly optimizes the camera poses and the radiance field by leveraging the asynchronous stream of events and calibrating sparse RGB frames. Bhattacharya et al. (2024) present an event-based dynamic NeRF, which is designed to faithfully reconstruct event streams in scenes with both rigid and non-rigid deformations, which are often too fast to capture with standard cameras. The approach improves test-time predictions of events at fine time resolutions by training on varied batch sizes of events. Qi et al. (2024a) propose a camera pose estimation framework to generalize the method to practical applications. They also introduce two novel losses: an event-enhanced blur rendering loss and an event rendering loss, which model the real blur and event generation processes, respectively. Feng et al. (2025) introduce a joint pose-NeRF training framework that includes a pose correction module to enhance the robustness of 3D reconstructions from inaccurate camera poses. They address the challenges of reconstructing NeRF from event data captured under non-ideal conditions, such as non-uniform event sequences, noisy poses and varying scene scales. Ev-NeRF (Hwang et al., 2023) leverages the multi-view consistency of NeRF to provide a self-supervision signal that filters out spurious measurements and extracts consistent underlying structures from the noisy event input. Instead of using posed images like traditional NeRF, Ev-NeRF uses event measurements along with sensor movements to create an integrated neural volume.

The additional prior derived from the event information helps to achieve a good photorealistic 3D reconstruction. Cannici and Scaramuzza (2024) combine model-based priors and learning-based modules to enhance NeRF reconstructions using event data and blurry frames. It explicitly models the blur formation process and uses the event double integral as an additional prior. Additionally, the method incorporates an end-to-end learnable response function to adapt to real event-camera sensor non-idealities. To overcome the challenges of existing methods that often rely on dense and low-noise event streams, Robust e-NeRF (Low and Lee, 2023) incorporates a realistic event generation model that accounts for intrinsic parameters and non-idealities, such as pixel-to-pixel threshold variations. Additionally, it uses a pair of normalized reconstruction losses that generalize effectively to arbitrary speed profiles and intrinsic parameter values without prior knowledge. Wang et al. (2024c) use motion, geometry and density priors to impose strong physical constraints, which significantly improve the robustness and efficiency of 3D scene reconstruction. This method leverages a density-guided patch-based sampling strategy, which accelerates the training process and enhances the expression of local geometries.

Advancement of 3DGS models with events. In contrast to NeRF’s implicit representation, 3DGS uses explicit 3D Gaussians to model scene geometry and appearance, enabling faster and often more interpretable reconstructions. Traditional 3DGS algorithms suffer in motion blur scenarios and depend heavily on sharp images and accurate camera poses. It is challenging to obtain in real-world scenarios with different real-world non-ideal settings. However, the incorporation of event information in 3DGS helps to improve the robustness of photorealistic 3D reconstruction and makes 3D reconstruction possible in different challenging scenarios. In this direction, Yura et al. (2025) use prior knowledge encoded in an E2V model for initializing the optimization process of 3DGS and use spline interpolation to obtain high-quality poses along the event camera trajectory. This method enhances reconstruction quality from fast-moving cameras while overcoming the computational limitations traditionally associated with event-based NeRF methods. Huang et al. (2025b) propose IncEventGS, an incremental 3DGS reconstruction algorithm using only a single event camera and without any information about camera pose. IncEventGS divides incoming event streams into chunks and models the camera trajectory as a continuous function. It randomly selects two close timestamps and integrates the corresponding event data. Using 3DGS, it renders two brightness images at the respective poses and minimizes the photometric loss between the synthesized and observed events. During initialization, a pre-trained depth estimation model infers depth from the rendered images to bootstrap the system. Weng et al. (2024b) use an adaptive deviation estimator (ADE) network to estimate Gaussian center deviations and use novel loss functions to achieve sharp 3D reconstructions in real time. In 3DGS, scenes are represented using Gaussian ellipsoids and each with a center position in 3D space. When input images are blurry due to camera motion, the initial positions of these Gaussians, which are estimated from blurry images, are inaccurate and this leads to poor 3D reconstructions. ADE network corrects these inaccuracies and estimates how much each Gaussian center should be deviated using the original Gaussian position and the camera poses estimated from sharp latent images, which are generated using event data. The novel loss function related to Gaussian center deviation is designed to simulate the motion blur process during exposure and guide the network to learn accurate Gaussian deviations. This approach demonstrates significant performance improvements over traditional 3DGS methods, particularly in scenarios with high-speed motion and low-light conditions. Yu et al. (2024b) model the formation process of motion-blurred images and guide the deblurring reconstruction of 3DGS by jointly optimizing 3DGS parameters and recovering camera motion trajectories during exposure. Event-3DGS (Han et al., 2025) framework processes event data directly and reconstructs 3D scenes by optimizing both scenario and sensor parameters. Key components include a high-pass filter-based photovoltage estimation module to reduce noise and an event-based 3D reconstruction loss to enhance reconstruction quality. Lee and Lee (2025) propose DiET-GS, which is a diffusion prior and event stream-driven MD in 3DGS. The diffusion prior leverages a pretrained diffusion model to guide the reconstruction toward more natural-looking images, compensating for artifacts introduced by event-based deblurring. DiET-GS addresses the concerns of inaccurate color and lost fine-grained details that are present in existing solutions. They constrain 3DGS with the event double integral to produce accurate color and details. They also leverage diffusion prior to enhancing the fine edge details.

Other possibilities in NeRF and 3DGS with events. Apart from already explored non-ideal situations like low-light and motion blur, where event information is useful, it can also be extended to other non-ideal capture scenarios like low-resolution capture, out-of-focus blur, depth-of-field limitations, multi-camera synchronization and multi-camera calibration. Events, due to their high temporal resolution, can effectively capture fine-grained changes in scene structure. When fused with intensity data or used to guide neural reconstruction models, they enable spatial super-resolution (Han et al., 2021; Kai et al., 2024, 2023; Liu et al., 2024a; Lu et al., 2023c; Xiao et al., 2024b; Yu et al., 2023), improving both texture fidelity and geometric precision. Out-of-focus blur and depth-of-field limitations are common in real-world intensity captures, particularly when using low-cost or wide-aperture lenses. In such cases, frame-based 3D reconstruction pipelines struggle to resolve scene details due to the loss of high-frequency information. Event cameras, being unaffected by traditional focus constraints, continue to respond to intensity changes regardless of depth-of-field issues (Bao et al., 2023; Lou et al., 2023; Teng et al., 2024). Precise calibration of intrinsic and extrinsic parameters is critical for reliable multi-view 3D reconstruction. In conventional systems, calibration relies on detecting specific patterns or correspondences across synchronized frames, which may fail under low resolution, low light, fast motion or in texture-less regions. Event cameras offer dense, temporally continuous observations of scene changes, which can be used to extract motion or edge-based features robustly across different views. This facilitates both offline and online calibration, improving the geometric consistency of reconstructed scenes. In summary, event-enhanced methods hold significant promise for making photorealistic 3D reconstruction more robust, scalable and adaptable to real-world conditions.

Event camera data set creation involves both real-world data collection and synthetic data generation techniques, each addressing different aspects of the data scarcity challenge in neuromorphic vision research. These approaches enable researchers to develop robust event-based algorithms despite the limited availability of physical event cameras. In this section, we present a curated list of event-based video data sets, highlighting their importance in advancing tasks such as video reconstruction, frame interpolation and enhancement under challenging conditions.

Data set creation methods for event camera data can be briefly divided into three categories based on the method used for creating the event stream along with the paired RGB data.

Synthetic data generation. Synthetically generating event stream data (Feng et al., 2025; Li et al., 2023; Yu et al., 2024b) offers the advantage of perfect and dense ground truth for tasks that are difficult to measure in the real world, such as optical flow, depth and 3D reconstruction. This process typically involves two stages. First, a high-frame-rate, photorealistic video sequence is rendered using a graphics engine like Blender. Second, this video is fed into an event camera simulator, which generates a corresponding event stream. This approach gives researchers full control over scene content, motion and lighting and provides flawless ground truth for camera poses, sharp frames and per-pixel information.

Real-world capture with beam-splitter rigs. For restoration tasks like MD, frame interpolation and low-light enhancement, achieving perfect spatial and temporal alignment between event streams and ground truth images is critical. The most common solution is a beam-splitter rig (Duan et al., 2025; Liu et al., 2024c, 2025b; Sun et al., 2023). This optical setup uses a semi-transparent mirror to direct incoming light into two identical optical paths – one leading to an event camera and the other to a high-speed, HR conventional camera. This ensures that both sensors capture the exact same scene from the same viewpoint at the same time, providing perfectly aligned, pixel-to-pixel ground truth.

Real-world capture with rigidly mounted sensors. Many data sets, especially those for robotics and autonomous driving applications like simultaneous localization and mapping and visual odometry, are captured using multi-sensor rigs (Gehrig et al., 2021; Hidalgo-Carrió et al., 2022; Klenk et al., 2021; Zhu et al., 2018). In this setup, an event camera is rigidly mounted alongside other sensors such as RGB cameras, inertial measurement units (IMUs) and sometimes LiDAR. The sensors are carefully calibrated to determine their geometric relationships (extrinsic parameters) and are synchronized using hardware triggers. While this method does not provide the pixel-perfect alignment of a beam splitter, it is ideal for capturing diverse, large-scale real-world scenarios.

Table 2 summarizes some of the most popular data sets by outlining the target tasks and key characteristics, including resolution, sensor modalities and duration. To provide additional clarity, we briefly categorize data sets based on their primary applications and data composition.

Table 2.

List of openly available data sets with event data stream

Data setTasksDescription
LED (Duan, 2024)Event denoising, low light enhancement, video reconstruction3000 sequences of high resolution (⁠1200×680⁠) paired data set
EventAID (Duan et al., 2025)Image and video enhancement, HDR conversion 1600+sec of capture with different texture and motion parameters
RLED (Liu et al., 2024c)Low light enhancement, frame interpolation, video reconstruction64200 aligned image and low-light event pairs
DSEC (Gehrig et al., 2021)Super-resolution, video reconstruction, occlusion removal53 sequences stereo RGB-Event camera pair along with LIDAR and GPS measurements
MVSEC (Zhu et al., 2018)HDR conversion, video reconstructionStereo capture of gray scale images, event stream and IMU readings
VECtor (Gao et al., 2022b)Occlusion removal, frame interpolationStereo capture of RGB event pair, depth sensor and IMU measurements
IJRR (Mueggler et al., 2017)Deblurring, frame interpolation24fps RG images with asynchronous event pair and IMU data
HQF (Stoffregen et al., 2020)Video reconstruction, super-resolutionEvent camera data set with two independent sensors
TemMat (Bao et al., 2024)Video reconstruction, frame interpolation, focus controlLow light high-dynamic range data set
HighREV (Sun et al., 2023)Frame interpolation, motion deblurringPaired event - RGB data
Erf-X170FPS (Kim et al., 2023)Frame interpolation, video interpolation, super-resolutionHigh resolution high-fps video frames paired with event data of extremely large motion scenes
CED (Scheerlinck et al., 2019)Colorization, super resolution, HDR conversionColor event data set contining 50 min of paired footage
BlinkFlow (Li et al., 2023)Frame interpolation, video reconstruction, motion deblurringSimulator for fast generation of event-based optical flow and associated data set
TUM-VIE (Klenk et al., 2021)3D reconstruction, video reconstruction, super resolutionStereo event camera data from a large variety of handheld and head-mounted sequences in indoor and outdoor environments
EDS (Hidalgo-Carrió et al., 2022)3D reconstruction, video, reconstruction, occlusion removalPaired event-RGB data along with IMU measurements
AE-NeRF (Feng et al., 2025)3D reconstruction, HDR conversion, occlusion removalSynthetic data set based on an improved version of ESIM
PAEv3d (Wang et al., 2024c)3D reconstruction, occlusion removal, video reconstructionLarge scale data set containing 101 obejcts along with depth maps
EvaGaussians (Yu et al., 2024b)Deblurring, video reconstruction, frame interpolationIndoor and outdoor synthetic scenes and six 3D objects created using blender, 5 real world scenes
Real-World-Blur (Qi et al., 2023)Deblurring, 3D reconstruction, video reconstructionColor event camera data set of 5 low-light blur scenes
Real-World-Challenge (Qi et al., 2023)Deblurring, 3D reconstruction, video reconstructionColor event camera data set of 5 low-light blur scenes
Dynamic EventNeRF (Rudnev et al., 2023)3D reconstruction, video reconstruction, super resolutionMulti-view event stream data set with 6 event cameras
CHMD (Liu et al., 2025b)Frame interpolationAligned data set of 90 sequences of high-speed non-linear moving targets
ClerMotion (Chen et al., 2025)Frame interpolationReal-world data set with object and camera motion

Despite the growing number of event-based video data sets, several limitations persist that hinder broader applicability and model generalization. Many data sets lack sufficient diversity in scenes, often focusing on constrained environments with limited variation in object types, motion patterns and backgrounds. Additionally, most data sets provide minimal or no semantic-level annotations, limiting their usefulness for tasks requiring high-level understanding, such as object recognition or scene parsing. There is also a noticeable scarcity of color event data sets, with the majority relying solely on grayscale event streams. Multi-view event data sets, which are crucial for tasks such as 3D reconstruction and view synthesis, remain underrepresented. Furthermore, the range of lighting conditions captured is often narrow, with few data sets systematically covering challenging scenarios such as extreme low light, HDR or rapid illumination changes. Addressing these gaps is essential for advancing event-based video research and enabling more robust and versatile models.

While this survey has extensively reviewed the contributions of event cameras to visual restoration and 3D reconstruction, an adjacent area of neuromorphic vision research centered on spike cameras is rapidly gaining momentum (Huang et al., 2023b). Sharing a bio-inspired origin, spike cameras also provide asynchronous, high-temporal-resolution data, but they operate on a fundamentally different principle: encoding absolute light intensity into a stream of binary spikes, as opposed to capturing brightness changes. This unique data modality has opened new avenues for tackling complex visual tasks, often complementing or offering alternative solutions to those provided by event cameras.

A spike camera operates on an “integrate-and-fire” model. Each pixel in a spike camera independently accumulates incoming photons (light). When the accumulated light intensity at a pixel reaches a predefined threshold, it fires a binary “spike” and then resets its accumulator. This process results in a continuous stream of spikes from the sensor. The key takeaway is that the frequency of spikes from a pixel is directly proportional to the intensity of light it receives. Brighter areas will produce more spikes over a given period than darker areas. This allows for a very fine-grained representation of light intensity over time.

Compared to event cameras, the fundamental difference lies in what triggers an output: while a change in brightness triggers output in event cameras, the intensity of light falling on the sensor causes the output trigger in spike cameras. We give a comparison of both camera sensors in Table 3.

Table 3.

Comparison of the spike camera sensor and the event camera sensor

FeatureSpike cameraEvent camera
TriggerAccumulation of light intensity reaching a thresholdChange in brightness (temporal contrast) exceeding a threshold
OutputA stream of binary spikesA stream of “events,” each containing the pixel coordinates, a timestamp and the polarity of the brightness change (positive or negative)
Information codedAbsolute brightness (encoded in the spike frequency)Change in brightness
Data representationA sequence of spikes over timeA sequence of asynchronous events

To support this growing field, the SpikeCV open-source framework (Zheng et al., 2023b) provides a unified toolbox with standardized data sets and algorithms. A primary focus of spike camera research has been overcoming motion-related challenges. Novel methods leverage spike streams to reconstruct clear, high-frame-rate temporal sequences from single, severely blurry real-world images (Chen et al., 2024b) and perform robust MD even when the spatiotemporal alignment between the spike and RGB data is unknown (Zhang et al., 2024c). Spike cameras have proven particularly effective for deblurring in extreme high-speed scenes (Chen et al., 2023a) and for enabling high-speed VFI where traditional methods fail (Xia et al., 2023). Beyond deblurring, hybrid spike-RGB camera systems are being used for complex reconstruction tasks, including high spatio-temporal video reconstruction (Xia et al., 2025) and generating sharp, all-in-focus images from a neuromorphic focal stack (Teng et al., 2024). Furthermore, these hybrid systems excel at HDR (HDR) imaging, fusing conventional images with spike data to create stunning HDR visuals (Han et al., 2023) and even producing 1000 FPS HDR video, pushing the boundaries of both temporal resolution and dynamic range simultaneously (Chang et al., 2023).

Current spike cameras are constrained by inherent hardware limitations, such as low spatial resolution and a lack of color information, which necessitate complex processing to fuse their data with conventional RGB sensors. Furthermore, practical deployment is challenged by the susceptibility of spike streams to noise and the strict requirement for precise spatiotemporal alignment in hybrid camera systems, where misalignment degrades quality. However, these very limitations in sensor fusion, algorithmic robustness and domain adaptation provide a rich landscape of compelling opportunities for future research.

Event cameras have emerged as a transformative sensing modality, offering high temporal resolution, low latency and HDR. This survey has comprehensively reviewed their integration with traditional frame-based systems for visual media restoration and 3D reconstruction, highlighting three key domains: temporal enhancement, spatial enhancement and 3D reconstruction. In the temporal domain, event cameras enable high-fidelity video reconstruction, frame interpolation and MD, especially under fast motion and low-light conditions. Their asynchronous nature allows for precise motion capture and temporal continuity, overcoming limitations of conventional cameras. In the spatial domain, event data contributes significantly to super-resolution, HDR imaging, low-light enhancement, occlusion removal and artifact reduction. By capturing fine-grained changes in brightness, event cameras complement RGB data to restore texture and detail in degraded scenes. In the 3D domain, event cameras facilitate dynamic and photorealistic 3D reconstruction, especially in challenging scenarios involving motion blur, low light and sparse data. Their integration with NeRF and 3DGS models has opened new avenues for real-time, high-quality scene modeling.

To further advance the field and guide future research, we propose the following research opportunities:

  • Event-driven multi-modal fusion: Develop robust fusion frameworks that dynamically balance RGB, depth and event modalities for real-time restoration and reconstruction.

  • Low-resource and edge deployment: Design lightweight architectures optimized for mobile and embedded platforms for event-based processing in resource-constrained environments.

  • Self-supervised and unsupervised earning: Explore domain adaptation, contrastive learning and generative models to reduce reliance on labeled data and improve generalization across scenes to solve the scarcity of event data.

  • Event-based calibration and synchronization: Investigate calibration-free multi-camera setups using event streams for robust synchronization and alignment in dynamic environments.

  • Event-guided generative models: Integrate event data with diffusion models for high-quality synthesis and restoration under extreme conditions.

  • Event-based depth and focus estimation: Advance depth-from-events and focus control algorithms for applications in robotics, augmented reality/virtual reality and computational photography.

  • Cross-domain applications: Apply event-based restoration techniques to domains such as medical imaging, remote sensing and industrial inspection.

  • Color event cameras: Expand research into color event data for improved visual media restoration and reconstruction.

  • Benchmarking and data set expansion: Create diverse, annotated and multi-view data sets covering varied lighting, motion and semantic contexts to support reproducible research.

  • Security and robustness: Address vulnerabilities in event-based systems, including adversarial attacks and backdoor threats, to ensure safe deployment in critical applications.

By consolidating recent progress and outlining future directions, this survey aims to serve as a foundational resource for researchers and practitioners seeking to harness the full potential of event cameras in visual media restoration and 3D reconstruction.

Afshar
,
S.
,
Hamilton
,
T.J.
,
Tapson
,
J.
,
Van Schaik
,
A.
and
Cohen
,
G.
(
2019
), “
Investigation of event-based surfaces for high-speed detection, unsupervised feature extraction, and object recognition
”,
Frontiers in Neuroscience
, Vol.
12
, p.
1047
.
Akman
,
A.
,
Kilicc
,
O.S.
and
Alatan
,
A.
(
2023
), “
MAEVI: motion aware event-based video frame interpolation
”, ArXiv, ,
available at:
Link to MAEVI: motion aware event-based video frame interpolationLink to the cited article.
Alzugaray
,
I.
and
Chli
,
M.
(
2018
), “
ACE: an efficient asynchronous corner tracker for event cameras
”,
2018 International Conference on 3D Vision (3DV)
,
IEEE
, pp.
653
-
661
.
Bai
,
W.
,
Chen
,
Y.
,
Feng
,
R.
and
Zheng
,
Y.
(
2022
), “
Accurate and efficient frame-based event representation for AER object recognition
”,
2022 International Joint Conference on Neural Networks (IJCNN)
,
IEEE
, pp.
1
-
6
.
Baldwin
,
R.W.
,
Liu
,
R.
,
Almatrafi
,
M.
,
Asari
,
V.
and
Hirakawa
,
K.
(
2022
), “
Time-ordered recent event (tore) volumes for event cameras
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
45
No.
2
, pp.
2519
-
2532
.
Bao
,
Y.
,
Gao
,
S.
,
Li
,
W.
and
Wang
,
K.
(
2025
), “
One-step event-driven high-speed autofocus
”,
Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR),
pp.
6222
-
6230
.
Bao
,
Y.
,
Sun
,
L.
,
Ma
,
Y.
and
Wang
,
K.
(
2024
), “
Temporal-mapping photography for event cameras
”,
European Conference on Computer Vision
,
Springer
, pp.
55
-
72
.
Bao
,
Y.
,
Sun
,
L.
,
Ma
,
Y.
,
Gu
,
D.
and
Wang
,
K.
(
2023
), “
Improving fast auto-focus with event polarity
”,
Optics Express
, Vol.
31
No.
15
, pp.
24025
-
24044
.
Bardow
,
P.
,
Davison
,
A.J.
and
Leutenegger
,
S.
(
2016
), “
Simultaneous optical flow and intensity estimation from an event camera
”,
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
, pp.
884
-
892
.
Baudron
,
A.
,
Wang
,
Z.W.
,
Cossairt
,
O.
and
Katsaggelos
,
A.K.
(
2020
), “
E3D: event-based 3d shape reconstruction
”,
arXiv preprint
.
Benosman
,
R.
,
Clercq
,
C.
,
Lagorce
,
X.
,
Ieng
,
S.H.
and
Bartolozzi
,
C.
(
2013
), “
Event-based visual flow
”,
IEEE Transactions on Neural Networks and Learning Systems
, Vol.
25
No.
2
, pp.
407
-
417
.
Bhattacharya
,
A.
,
Madaan
,
R.
,
Cladera
,
F.
,
Vemprala
,
S.
,
Bonatti
,
R.
,
Daniilidis
,
K.
,
Kapoor
,
A.
,
Kumar
,
V.
,
Matni
,
N.
and
Gupta
,
J.K.
(
2024
), “
EVDNERF: reconstructing event data with dynamic neural radiance fields
”,
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision
, pp.
5846
-
5855
.
Bi
,
Y.
,
Chadha
,
A.
,
Abbas
,
A.
,
Bourtsoulatze
,
E.
and
Andreopoulos
,
Y.
(
2020
), “
Graph-based spatio-temporal feature learning for neuromorphic vision sensing
”,
IEEE Transactions on Image Processing
, Vol.
29
, pp.
9084
-
9098
.
Brandli
,
C.
,
Berner
,
R.
,
Yang
,
M.
,
Liu
,
S.C.
and
Delbruck
,
T.
(
2014
), “
A 240× 180 130 db 3 μs latency global shutter spatiotemporal vision sensor
”,
IEEE Journal of Solid-State Circuits
, Vol.
49
No.
10
, pp.
2333
-
2341
.
Cadena
,
P.R.G.
,
Qian
,
Y.
,
Wang
,
C.
and
Yang
,
M.
(
2021
), “
SPADEE2VID: spatially-adaptive denormalization for event-based video reconstruction
”,
IEEE Transactions on Image Processing: a Publication of the IEEE Signal Processing Society
, Vol.
30
, pp.
2488
-
2500
,
available at:
Link to SPADEE2VID: spatially-adaptive denormalization for event-based video reconstructionLink to the cited article.
Cannici
,
M.
and
Scaramuzza
,
D.
(
2024
), “
Mitigating motion blur in neural radiance fields with events and frames
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
9286
-
9296
.
Cannici
,
M.
,
Ciccone
,
M.
,
Romanoni
,
A.
and
Matteucci
,
M.
(
2020
), “
A differentiable recurrent surface for asynchronous event-based data
”,
Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16
,
Springer
, pp.
136
-
152
.
Cao
,
C.
,
Fu
,
X.
,
Zhu
,
Y.
,
Sun
,
Z.
and
Zha
,
Z.J.
(
2023
), “
Event-driven video restoration with spiking-convolutional architecture
”,
IEEE Transactions on Neural Networks and Learning Systems
, Vol.
36
No.
1
.
Cao
,
R.
,
Galor
,
D.
,
Kohli
,
A.
,
Yates
,
J.L.
and
Waller
,
L.
(
2024
), “
Noise2Image: noise-enabled static scene recovery for event cameras
”, ArXiv, ,
available at:
Link to Noise2Image: noise-enabled static scene recovery for event camerasLink to a PDF of the cited article.
Carneiro
,
J.
,
Ieng
,
S.H.
,
Posch
,
C.
and
Benosman
,
R.
(
2013
), “
Eventbased 3D reconstruction from neuromorphic retinas
”,
Neural Networks: The Official Journal of the International Neural Network Society
, Vol.
45
, pp.
27
-
38
.
Chakravarthi
,
B.
,
Verma
,
A.A.
,
Daniilidis
,
K.
,
Fermuller
,
C.
and
Yang
,
Y.
(
2024
), “
Recent event camera innovations: a survey
”,
arXiv preprint
.
Chang
,
Y.
,
Zhou
,
C.
,
Hong
,
Y.
,
Hu
,
L.
,
Xu
,
C.
,
Huang
,
T.
and
Shi
,
B.
(
2023
), “
1000 Fps HDR video with a spike-RGB hybrid camera
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
22180
-
22190
.
Chen
,
J.
,
Feng
,
B.Y.
,
Cai
,
H.
,
Wang
,
T.
,
Burner
,
L.
,
Yuan
,
D.
,
Fermuller
,
C.
,
Metzler
,
C.A.
and
Aloimonos
,
Y.
(
2025
), “
Repurposing pre-trained video diffusion models for event-based video interpolation
”,
Proceedings of the Computer Vision and Pattern Recognition Conference
, pp.
12456
-
12466
.
Chen
,
J.
,
Zhu
,
Y.
,
Lian
,
D.
,
Yang
,
J.
,
Wang
,
Y.
,
Zhang
,
R.
,
Liu
,
X.
,
Qian
,
S.
,
Kneip
,
L.
and
Gao
,
S.
(
2023b
), “
Revisiting event-based video frame interpolation
”,
2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
,
IEEE
, pp.
1292
-
1299
.
Chen
,
K.
,
Chen
,
S.
,
Zhang
,
J.
,
Zhang
,
B.
,
Zheng
,
Y.
,
Huang
,
T.
and
Yu
,
Z.
(
2024b
), “
Spikereveal: unlocking temporal sequences from real blurry inputs with spike streams
”,
Advances in Neural Information Processing Systems
, Vol.
37
, pp.
62673
-
62696
.
Chen
,
K.
,
Liang
,
G.
,
Li
,
H.
,
Lu
,
Y.
and
Wang
,
L.
(
2024a
), “
EvLight++: low-light video enhancement with an event camera: a large-scale real-world data set, novel method, and more
”,
arXiv preprint
.
Chen
,
S.
,
Zhang
,
J.
,
Zheng
,
Y.
,
Huang
,
T.
and
Yu
,
Z.
(
2023a
), “
Enhancing motion deblurring in high-speed scenes with spike streams
”,
Advances in Neural Information Processing Systems
, Vol.
36
, pp.
70279
-
70292
.
Chen
,
T.
,
Kornblith
,
S.
,
Norouzi
,
M.
and
Hinton
,
G.
(
2020
), “
A simple framework for contrastive learning of visual representations
”,
International Conference on Machine Learning
,
PmLR
, pp.
1597
-
1607
.
Chen
,
Y.
,
Liu
,
S.
and
Wang
,
X.
(
2021
), “
Learning continuous image representation with local implicit image function
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
8628
-
8638
.
Cheng
,
L.
,
Liu
,
N.
,
Guo
,
X.
,
Shen
,
Y.
,
Meng
,
Z.
,
Huang
,
K.
and
Zhang
,
X.
(
2022
), “
A novel rain removal approach for outdoor dynamic vision sensor event videos
”,
Frontiers in Neurorobotics
, Vol.
16
, p.
928707
.
Cheng
,
Z.
,
Zhang
,
X.
,
Yu
,
L.
,
Liu
,
J.Z.
,
Yang
,
W.
and
Xia
,
G.S.
(
2023
), “
Recovering continuous scene dynamics from a single blurry image with events
”, ArXiv, ,
available at:
Link to Recovering continuous scene dynamics from a single blurry image with eventsLink to the cited article.
Cho
,
H.
,
Jeong
,
Y.
,
Kim
,
T.
and
Yoon
,
K.J.
(
2023
), “
Non-coaxial event guided motion deblurring with spatial alignment
”,
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
, pp.
12458
-
12469
,
available at:
Link to Non-coaxial event guided motion deblurring with spatial alignmentLink to the cited article.
Cho
,
H.
,
Kim
,
T.
,
Jeong
,
Y.
and
Yoon
,
K.J.
(
2024
), “
TTA-EVF: test-time adaptation for event-based video frame interpolation via reliable pixel and sample estimation
”,
2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
25701
-
25711
,
available at:
Link to TTA-EVF: test-time adaptation for event-based video frame interpolation via reliable pixel and sample estimationLink to a PDF of the cited article.
Choudhury
,
A.
and
Su
,
G.M.
(
2025
), “
Modeling event camera frame sequence using neural field
”,
IEEE International Conference on Multimedia Information Processing and Retrieval
.
Choudhury
,
A.
,
Su
,
G.M.
and
Chen
,
J.
(
2025
), “
Triplane learning for event stream representation
”,
59th Asilomar Conference on Signals, Systems and Computers
.
Cook
,
M.
,
Gugelmann
,
L.
,
Jug
,
F.
,
Krautz
,
C.
and
Steger
,
A.
(
2011
), “
Interacting maps for fast visual interpretation
”,
The 2011 International Joint Conference on Neural Networks
,
IEEE
, pp.
770
-
776
,
Cui
,
M.
,
Wang
,
Z.
,
Wang
,
D.
,
Zhao
,
B.
and
Li
,
X.
(
2024
), “
Color event enhanced single-exposure HDR imaging
”,
Proceedings of the AAAI Conference on Artificial Intelligence
, Vol.
38
No.
2
, pp.
1399
-
1407
.
Deng
,
Y.
,
Chen
,
H.
,
Liu
,
H.
and
Li
,
Y.
(
2022
), “
A voxel graph CNN for object classification with event cameras
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
1172
-
1181
.
Deng
,
Y.
,
Li
,
Y.
and
Chen
,
H.
(
2020
), “
AMAE: adaptive motion-agnostic encoder for event-based object classification
”,
IEEE Robotics and Automation Letters
, Vol.
5
No.
3
, pp.
4596
-
4603
.
Ding
,
S.
,
Chen
,
J.
,
Wang
,
Y.
,
Kang
,
Y.
,
Song
,
W.
,
Cheng
,
J.
and
Cao
,
Y.
(
2023
), “
E-MLB: multilevel benchmark for event-based camera denoising
”,
IEEE Transactions on Multimedia
, Vol.
26
, pp.
65
-
76
.
Duan
,
P.
,
Li
,
B.
,
Yang
,
Y.
,
Lou
,
H.
,
Teng
,
M.
,
Zhou
,
X.
,
Ma
,
Y.
and
Shi
,
B.
(
2025
), “
EventAid: benchmarking event-aided image/video enhancement algorithms with real-captured hybrid data set
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
47
No.
8
.
Duan
,
P.
,
Ma
,
Y.
,
Zhou
,
X.
,
Shi
,
X.
,
Wang
,
Z.W.
,
Huang
,
T.
and
Shi
,
B.
(
2023
), “
Neurozoom: denoising and super resolving neuromorphic events and spikes
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
45
No.
12
, pp.
15219
-
15232
.
Duan
,
P.
,
Wang
,
Z.W.
,
Zhou
,
X.
,
Ma
,
Y.
and
Shi
,
B.
(
2021
), “
EventZoom: learning to denoise and super resolve neuromorphic events
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
12824
-
12833
.
Duan
,
Y.
(
2024
), “
LED: a large-scale real-world paired data set for event camera denoising
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
25637
-
25647
.
Ercan
,
B.
,
Eker
,
O.
,
Saglam
,
C.
,
Erdem
,
A.
and
Erdem
,
E.
(
2024
), “
Hypere2VID: improving event-based video reconstruction via hyper networks
”,
IEEE Transactions on Image Processing
, Vol.
33
.
Feng
,
C.
,
Yu
,
W.
,
Cheng
,
X.
,
Tang
,
Z.
,
Zhang
,
J.
,
Yuan
,
L.
and
Tian
,
Y.
(
2025
), “
AE-NeRF: augmenting event-based neural radiance fields for non-ideal conditions and larger scene
”,
arXiv preprint
.
Fu
,
X.
,
Cao
,
C.
,
Xu
,
S.
,
Zhang
,
F.
,
Wang
,
K.
and
Zha
,
Z.J.
(
2024
), “
Event-driven heterogeneous network for video deraining
”,
International Journal of Computer Vision
, Vol.
132
No.
12
, pp.
5841
-
5861
.
Gallego
,
G.
,
Delbrück
,
T.
,
Orchard
,
G.
,
Bartolozzi
,
C.
,
Taba
,
B.
,
Censi
,
A.
,
Leutenegger
,
S.
,
Davison
,
A.J.
,
Conradt
,
J.
,
Daniilidis
,
K.
and
Scaramuzza
,
D.
(
2020
), “
Event-based vision: a survey
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
44
No.
1
, pp.
154
-
180
.
Gallego
,
G.
,
Gehrig
,
M.
and
Scaramuzza
,
D.
(
2019
), “
Focus is all you need: loss functions for event-based vision
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
12280
-
12289
.
Gao
,
L.
,
Liang
,
Y.
,
Yang
,
J.
,
Wu
,
S.
,
Wang
,
C.
,
Chen
,
J.
and
Kneip
,
L.
(
2022b
), “
VECtor: a versatile event-centric benchmark for multi-sensor SLAM
”,
IEEE Robotics and Automation Letters
, Vol.
7
No.
3
, pp.
8217
-
8224
.
Gao
,
T.
,
Yao
,
X.
and
Chen
,
D.
(
2021
), “
Simcse: simple contrastive learning of sentence embeddings
”,
arXiv preprint
.
Gao
,
Y.
,
Li
,
S.
,
Li
,
Y.
,
Guo
,
Y.
and
Dai
,
Q.
(
2022a
), “
SuperFast: 200× video frame interpolation via event camera
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
45
No.
6
, pp.
7764
-
7780
,
available at:
Link to SuperFast: 200× video frame interpolation via event cameraLink to the cited article.
Ge
,
C.
,
Fu
,
X.
,
He
,
P.
,
Wang
,
K.
,
Cao
,
C.
and
Zha
,
Z.J.
(
2024
), “
Neuromorphic event signal-driven network for video de-raining
”,
Proceedings of the AAAI Conference on Artificial Intelligence
, Vol.
38
No.
3
, pp.
1878
-
1886
.
Ge
,
C.
,
Fu
,
X.
,
He
,
P.
,
Wang
,
K.
,
Cao
,
C.
and
Zha
,
Z.J.
(
2025
), “
Event-Mamba: enhancing spatio-temporal locality with state space models for event-based video reconstruction
”,
Proceedings of the AAAI Conference on Artificial Intelligence
, Vol.
39
No.
3
, pp.
3104
-
3112
.
Gehrig
,
D.
and
Scaramuzza
,
D.
(
2022
), “
Are high-resolution event cameras really needed?
”,
arXiv preprint
.
Gehrig
,
D.
,
Loquercio
,
A.
,
Derpanis
,
K.G.
and
Scaramuzza
,
D.
(
2019
), “
End-to-end learning of representations for asynchronous event based data
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
5633
-
5643
.
Gehrig
,
M.
,
Aarents
,
W.
,
Gehrig
,
D.
and
Scaramuzza
,
D.
(
2021
), “
Dsec: a stereo event camera data set for driving scenarios
”,
IEEE Robotics and Automation Letters
, Vol.
6
No.
3
, pp.
4947
-
4954
.
Gu
,
A.
and
Dao
,
T.
(
2023
), “
Mamba: linear-time sequence modeling with selective state spaces
”,
arXiv preprint
.
Gu
,
A.
,
Goel
,
K.
and
Ré
,
C.
(
2021a
), “
Efficiently modeling long sequences with structured state spaces
”,
arXiv preprint
.
Gu
,
C.
,
Learned-Miller
,
E.
,
Sheldon
,
D.
,
Gallego
,
G.
and
Bideau
,
P.
(
2021c
), “
The spatio-temporal Poisson point process: a simple model for the alignment of event camera data
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
13495
-
13504
.
Gu
,
D.
,
Li
,
J.
,
Zhang
,
Y.
and
Tian
,
Y.
(
2021b
), “
How to learn a domain adaptive event simulator?
”,
Proceedings of the 29th ACM International Conference on Multimedia
, pp.
1275
-
1283
.
Gu
,
F.
,
Sng
,
W.
,
Taunyazov
,
T.
and
Soh
,
H.
(
2020
), “
Tactilesgnet: a spiking graph neural network for event-based tactile object recognition
”,
2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
,
IEEE
, pp.
9876
-
9882
.
Guo
,
G.
,
Feng
,
Y.
,
Lv
,
H.
,
Zhao
,
Y.
,
Liu
,
H.
and
Bi
,
G.
(
2023
), “
Eventguided image super-resolution reconstruction
”,
Sensors
, Vol.
23
No.
4
, p.
2155
.
Guo
,
Q.
,
Shi
,
H.
,
Li
,
H.
,
Xiao
,
J.
and
Gao
,
X.
(
2024a
), “
REDIR: refocus-free event-based de-occlusion image reconstruction
”,
European Conference on Computer Vision
,
Springer
, pp.
419
-
435
.
Guo
,
S.
,
Chen
,
Z.
,
Zhang
,
Z.
,
Chen
,
Y.
,
Xu
,
G.
and
Xue
,
T.
(
2024c
), “
Event-assisted 12-stop HDR imaging of dynamic scene
”,
arXiv preprint
.
Guo
,
Y.
,
Peng
,
W.
,
Chen
,
Y.
,
Zhou
,
J.
and
Ma
,
Z.
(
2024b
), “
Improved event-based image de-occlusion
”,
IEEE Signal Processing Letters
, Vol.
31
.
Han
,
H.
,
Li
,
J.
,
Wei
,
H.
and
Ji
,
X.
(
2025
), “
Event-3DGS: event-based 3D reconstruction using 3D Gaussian splatting
”,
Advances in Neural Information Processing Systems
, Vol.
37
, pp.
128139
-
128159
.
Han
,
H.
,
Lyu
,
J.
,
Li
,
J.
,
Wei
,
H.
,
Li
,
C.
,
Wei
,
Y.
,
Chen
,
S.
and
Ji
,
X.
(
2024
), “
Physical-based event camera simulator
”,
European Conference on Computer Vision
,
Springer
, pp.
19
-
35
.
Han
,
J.
,
Yang
,
Y.
,
Duan
,
P.
,
Zhou
,
C.
,
Ma
,
L.
,
Xu
,
C.
,
Huang
,
T.
,
Sato
,
I.
and
Shi
,
B.
(
2023
), “
Hybrid high dynamic range imaging fusing neuromorphic and conventional images
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
45
no No.
7
, pp.
8553
-
8565
.
Han
,
J.
,
Yang
,
Y.
,
Zhou
,
C.
,
Xu
,
C.
and
Shi
,
B.
(
2021
), “
EVINTSR-net: event guided multiple latent frames reconstruction and super resolution
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
4882
-
4891
.
Han
,
J.
,
Zhou
,
C.
,
Duan
,
P.
,
Tang
,
Y.
,
Xu
,
C.
,
Xu
,
C.
,
Huang
,
T.
and
Shi
,
B.
(
2020
), “
Neuromorphic camera guided high dynamic range imaging
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
1730
-
1739
.
He
,
K.
,
Fan
,
H.
,
Wu
,
Y.
,
Xie
,
S.
and
Girshick
,
R.
(
2020
), “
Momentum contrast for unsupervised visual representation learning
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
9729
-
9738
.
He
,
W.
,
You
,
K.
,
Qiao
,
Z.
,
Jia
,
X.
,
Zhang
,
Z.
,
Wang
,
W.
,
Lu
,
H.
,
Wang
,
Y.
and
Liao
,
J.
(
2022
), “
Timereplayer: unlocking the potential of event cameras for video interpolation
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
17804
-
17813
.
Hidalgo-Carrió
,
J.
,
Gallego
,
G.
and
Scaramuzza
,
D.
(
2022
), “
Event-aided direct sparse odometry
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
5781
-
5790
.
Ho
,
J.
,
Jain
,
A.
and
Abbeel
,
P.
(
2020
), “
Denoising diffusion probabilistic models
”, ArXiv, ,
available at:
Link to Denoising diffusion probabilistic modelsLink to the cited article.
Hu
,
Y.
,
Delbruck
,
T.
and
Liu
,
S.C.
(
2020
), “
Learning to exploit multiple vision modalities by using grafted networks
”,
European Conference on Computer Vision
,
Springer
, pp.
85
-
101
.
Hu
,
Y.
,
Liu
,
S.C.
and
Delbruck
,
T.
(
2021
), “
v2e: from video frames to realistic DVS events
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
1312
-
1321
.
Huang
,
J.
,
Dong
,
C.
,
Chen
,
X.
and
Liu
,
P.
(
2025b
), “
IncEventGS: pose-free Gaussian splatting from a single event camera
”,
Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR),
pp.
26933
-
26942
.
Huang
,
T.
,
Zheng
,
Y.
,
Yu
,
Z.
,
Chen
,
R.
,
Li
,
Y.
,
Xiong
,
R.
,
Ma
,
L.
,
Zhao
,
J.
,
Dong
,
S.
,
Zhu
,
L.
,
Li
,
J.
,
Jia
,
S.
,
Fu
,
Y.
,
Shi
,
B.
,
Wu
,
S.
and
Tian
,
Y.
(
2023b
), “
1000× Faster camera and machine vision with ordinary devices
”,
Engineering
, Vol.
25
, pp.
110
-
119
.
Huang
,
T.W.
,
Choudhury
,
A.
and
Su
,
G.M.
(
2025a
), “
Neural representations for event voxel grid
”,
IEEE International Conference on Image Processing (ICIP) 2025 Workshop on Generative AI for World Simulations and Communications
.
Huang
,
Z.
,
Liang
,
Q.
,
Yu
,
Y.
,
Qin
,
C.
,
Zheng
,
X.
,
Huang
,
K.
,
Zhou
,
Z.
and
Yang
,
W.
(
2024
), “
Bilateral event mining and complementary for event stream super-resolution
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
34
-
43
.
Huang
,
Z.
,
Sun
,
L.
,
Zhao
,
C.
,
Li
,
S.
and
Su
,
S.
(
2023a
), “
Eventpoint: Self supervised interest point detection and description for event based camera
”,
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision
, pp.
5396
-
5405
.
Hwang
,
I.
,
Kim
,
J.
and
Kim
,
Y.M.
(
2023
), “
EV-nerf: event based neural radiance field
”,
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision
, pp.
837
-
847
.
Jian
,
D.
and
Rostami
,
M.
(
2023
), “
Unsupervised domain adaptation for training event-based networks using contrastive learning and uncorrelated conditioning
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
18721
-
18731
.
Jiang
,
B.
,
Xiong
,
B.
,
Qu
,
B.
,
Salman Asif
,
M.
,
Zhou
,
Y.
and
Ma
,
Z.
(
2024a
), “
EDformer: transformer-based event denoising across varied noise levels
”,
European Conference on Computer Vision
,
Springer
, pp.
200
-
216
.
Jiang
,
J.
,
Zhou
,
X.
,
Wang
,
B.
,
Deng
,
X.
,
Xu
,
C.
and
Shi
,
B.
(
2024b
), “
Complementing event streams and RGB frames for hand mesh reconstruction
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
24944
-
24954
.
Jiang
,
Y.
,
Wang
,
Y.
,
Li
,
S.
,
Zhang
,
Y.
,
Zhao
,
M.
and
Gao
,
Y.
(
2023
), “
Event-based low-illumination image enhancement
”,
IEEE Transactions on Multimedia
, Vol.
26
, pp.
1920
-
1931
.
Jin
,
X.
,
Wu
,
L.
,
Chen
,
J.
,
Chen
,
Y.
,
Koo
,
J.
and
Hahm
,
C.H.
(
2022
), “
A unified pyramid recurrent network for video frame interpolation
”,
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
1578
-
1587
,
available at:
Link to A unified pyramid recurrent network for video frame interpolationLink to the cited article.
Jing
,
Y.
,
Yang
,
Y.
,
Wang
,
X.
,
Song
,
M.
and
Tao
,
D.
(
2021
), “
Turning frequency to resolution: video super-resolution via event cameras
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
7772
-
7781
.
Kai
,
D.
,
Lu
,
J.
,
Zhang
,
Y.
and
Sun
,
X.
(
2024
), “
EvTexture: event-driven texture enhancement for video super-resolution
”,
Forty-first International Conference on Machine Learning
.
Kai
,
D.
,
Zhang
,
Y.
and
Sun
,
X.
(
2023
), “
Video super-resolution via event driven temporal alignment
”,
2023 IEEE International Conference on Image Processing (ICIP)
,
IEEE
, pp.
2950
-
2954
.
Khosla
,
P.
,
Teterwak
,
P.
,
Wang
,
C.
,
Sarna
,
A.
,
Tian
,
Y.
,
Isola
,
P.
,
Maschinot
,
A.
,
Liu
,
C.
and
Krishnan
,
D.
(
2020
), “
Supervised contrastive learning
”,
Advances in Neural Information Processing Systems
, Vol.
33
, pp.
18661
-
18673
.
Kilicc
,
O.S.
,
Akman
,
A.
and
Alatan
,
A.
(
2022
), “
E-VFIA: event-based video frame interpolation with attention
”,
2023 IEEE International Conference on Robotics and Automation (ICRA)
, pp.
8284
-
8290
.
Kim
,
H.
,
Handa
,
A.
,
Benosman
,
R.
,
Ieng
,
S.H.
and
Davison
,
A.J.
(
2008
), “
Simultaneous mosaicing and tracking with an event camera
”,
J. Solid State Circ
, Vol.
43
, pp.
566
-
576
.
Kim
,
H.
,
Leutenegger
,
S.
and
Davison
,
A.J.
(
2016
), “
Real-time 3D reconstruction and 6-DoF tracking with an event camera
”,
European Conference on Computer Vision
,
available at:
Link to Real-time 3D reconstruction and 6-DoF tracking with an event cameraLink to the cited article.
Kim
,
T.
,
Chae
,
Y.
,
Jang
,
H.K.
and
Yoon
,
K.J.
(
2023
), “
Event-based video frame interpolation with cross-modal asymmetric bidirectional motion fields
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
18032
-
18042
.
Kim
,
T.
,
Cho
,
H.
and
Yoon
,
K.J.
(
2024a
), “
CMTA: cross-modal temporal alignment for event-guided video deblurring
”, ArXiv, ,
available at:
Link to CMTA: cross-modal temporal alignment for event-guided video deblurringLink to the cited article.
Kim
,
T.
,
Jeong
,
J.
,
Cho
,
H.
,
Jeong
,
Y.
and
Yoon
,
K.J.
(
2024b
), “
Towards real-world event-guided low-light video enhancement and deblurring
”,
European Conference on Computer Vision
,
Springer
, pp.
433
-
451
.
Kim
,
T.
,
Lee
,
J.
,
Wang
,
L.
and
Yoon
,
K.J.
(
2021
), “
Event-guided deblurring of unknown exposure time videos
”, ArXiv, ,
available at:
Link to Event-guided deblurring of unknown exposure time videosLink to the cited article.
Klenk
,
S.
,
Bonello
,
D.
,
Koestler
,
L.
,
Araslanov
,
N.
and
Cremers
,
D.
(
2024
), “
Masked event modeling: self-supervised pretraining for event cameras
”,
Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision
, pp.
2378
-
2388
.
Klenk
,
S.
,
Chui
,
J.
,
Demmel
,
N.
and
Cremers
,
D.
(
2021
), “
TUM-VIE: the TUM stereo visual-inertial event data set
”,
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
,
IEEE
, pp.
8601
-
8608
.
Klenk
,
S.
,
Koestler
,
L.
,
Scaramuzza
,
D.
and
Cremers
,
D.
(
2023
), “
ENERF: neural radiance fields from a moving event camera
”,
IEEE Robotics and Automation Letters
, Vol.
8
No.
3
, pp.
1587
-
1594
.
Lagorce
,
X.
,
Orchard
,
G.
,
Galluppi
,
F.
,
Shi
,
B.E.
and
Benosman
,
R.B.
(
2016
), “
Hots: a hierarchy of event-based time-surfaces for pattern recognition
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
39
no No.
7
, pp.
1346
-
1359
.
Lee
,
C.
,
Kosta
,
A.K.
,
Zhu
,
A.Z.
,
Chaney
,
K.
,
Daniilidis
,
K.
and
Roy
,
K.
(
2020
), “
Spike-flownet: event-based optical flow estimation with energy-efficient hybrid neural networks
”,
European Conference on Computer Vision
,
Springer
, pp.
366
-
382
.
Lee
,
S.
and
Lee
,
G.H.
(
2025
), “
DiET-GS: diffusion prior and event stream-assisted motion deblurring 3D Gaussian splatting
”,
Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR)
, pp.
21739
-
21749
.
Lehtinen
,
J.
,
Munkberg
,
J.
,
Hasselgren
,
J.
,
Laine
,
S.
,
Karras
,
T.
,
Aittala
,
M.
and
Aila
,
T.
(
2018
), “
Noise2Noise: learning image restoration without clean data
”,
International Conference on Machine Learning
,
PMLR
, pp.
2965
-
2974
.
Lei
,
T.
,
Guo
,
X.
and
Li
,
Y.
(
2024
), “
How many events are needed for one reconstructed image using an event camera?
”,
The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences
, Vol. XLVIII-4-2024, pp.
645
-
650
.
Li
,
S.
,
Feng
,
Y.
,
Li
,
Y.
,
Jiang
,
Y.
,
Zou
,
C.
and
Gao
,
Y.
(
2021
), “
Event stream super-resolution via spatiotemporal constraint learning
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
4480
-
4489
.
Li
,
S.Q.
,
Gao
,
Y.
and
Dai
,
Q.H.
(
2022
), “
Image de-occlusion via event enhanced multi-modal fusion hybrid network
”,
Machine Intelligence Research
, Vol.
19
No.
4
, pp.
307
-
318
.
Li
,
W.
,
Wan
,
P.
,
Wang
,
P.
,
Li
,
J.
,
Zhou
,
Y.
and
Liu
,
P.
(
2024c
), “
BeNeRF: neural radiance fields from a single blurry image and event stream
”,
European Conference on Computer Vision
,
Springer
, pp.
416
-
434
.
Li
,
X.
,
Cheng
,
S.
,
Zeng
,
Z.
,
Zhao
,
C.
and
Fan
,
C.
(
2024a
), “
ERS-HDRI: event-based remote sensing HDR imaging
”,
Remote Sensing
, Vol.
16
No.
3
, p.
437
.
Li
,
X.
,
Lu
,
Q.
,
Fan
,
C.
,
Zhao
,
C.
,
Zou
,
L.
and
Yu
,
L.
(
2024b
), “
Generalizing event-based HDR imaging to various exposures
”,
Neurocomputing
, Vol.
600
, p.
128132
.
Li
,
Y.
,
Huang
,
Z.
,
Chen
,
S.
,
Shi
,
X.
,
Li
,
H.
,
Bao
,
H.
,
Cui
,
Z.
and
Zhang
,
G.
(
2023
), “
Blinkflow: a data set to push the limits of event based optical flow estimation
”,
2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
,
IEEE
, pp.
3881
-
3888
.
Liang
,
G.
,
Chen
,
K.
,
Li
,
H.
,
Lu
,
Y.
and
Wang
,
L.
(
2024a
), “
Towards robust event-guided low-light image enhancement: a large-scale real world event-image data set and novel approach
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
23
-
33
.
Liang
,
J.
,
Yang
,
Y.
,
Li
,
B.
,
Duan
,
P.
,
Xu
,
Y.
and
Shi
,
B.
(
2023
), “
Coherent event guided low-light video enhancement
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
10615
-
10625
.
Liang
,
J.
,
Yu
,
B.
,
Yang
,
Y.
,
Han
,
Y.
and
Shi
,
B.
(
2024b
), “
E2VIDiff: perceptual events-to-video reconstruction using diffusion priors
”, ArXiv, ,
available at:
Link to E2VIDiff: perceptual events-to-video reconstruction using diffusion priorsLink to the cited article.
Liang
,
Q.
,
Huang
,
Z.
,
Zheng
,
X.
,
Yang
,
F.
,
Peng
,
J.
,
Huang
,
K.
and
Tian
,
Y.
(
2024c
), “
Efficient event stream super-resolution with recursive multi-branch fusion
”,
arXiv preprint
.
Liao
,
W.
,
Zhang
,
X.
,
Yu
,
L.
,
Lin
,
S.
,
Yang
,
W.
and
Qiao
,
N.
(
2022
), “
Synthetic aperture imaging with events and frames
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
17735
-
17744
.
Lin
,
M.
,
Zhang
,
C.
,
He
,
C.
and
Yu
,
L.
(
2023a
), “
Learning parallax for stereo event-based motion deblurring
”, ArXiv, ,
available at:
Link to Learning parallax for stereo event-based motion deblurringLink to the cited article.
Lin
,
S.
,
Ma
,
Y.
,
Guo
,
Z.
and
Wen
,
B.
(
2022a
), “
DVS-voltmeter: stochastic process-based event simulator for dynamic vision sensors
”,
European Conference on Computer Vision
,
Springer
, pp.
578
-
593
.
Lin
,
S.
,
Zhang
,
J.
,
Pan
,
J.
,
Jiang
,
Z.
,
Zou
,
D.
,
Wang
,
Y.
,
Chen
,
J.
and
Ren
,
J.
(
2020
), “
Learning event-driven video deblurring and interpolation
”,
Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16
,
Springer
, pp.
695
-
710
.
Lin
,
S.
,
Zhang
,
Y.
,
Yu
,
L.
,
Zhou
,
B.
,
Luo
,
X.
and
Pan
,
J.
(
2022b
), “
Autofocus for event cameras
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
16344
-
16353
.
Lin
,
X.
,
Qiu
,
C.
,
Shen
,
S.
,
Zang
,
Y.
,
Liu
,
W.
,
Bian
,
X.
,
Müller
,
M.
,
Wang
,
C.
and
Cai
,
Z.
(
2023b
), “
E2pnet: event to point cloud registration with spatio-temporal representation learning
”,
Advances in Neural Information Processing Systems
, Vol.
36
, pp.
18076
-
18089
.
Liu
,
H.
,
Peng
,
S.
,
Zhu
,
L.
,
Chang
,
Y.
,
Zhou
,
H.
and
Yan
,
L.
(
2024c
), “
Seeing motion at night time with an event camera
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
25648
-
25658
.
Liu
,
H.
,
Xu
,
J.
,
Chang
,
Y.
,
Zhou
,
H.
,
Zhao
,
H.
,
Wang
,
L.
and
Yan
,
L.
(
2025b
), “
TimeTracker: event-based continuous point tracking for video frame interpolation with non-linear motion
”,
Proceedings of the Computer Vision and Pattern Recognition Conference
, pp.
17649
-
17659
.
Liu
,
L.
,
An
,
J.
,
Liu
,
J.
,
Yuan
,
S.
,
Chen
,
X.
,
Zhou
,
W.
,
Li
,
H.
,
Wang
,
Y.F.
and
Tian
,
Q.
(
2023
), “
Low-light video enhancement with synthetic event guidance
”,
Proceedings of the AAAI Conference on Artificial Intelligence
, Vol.
37
No.
2
, pp.
1692
-
1700
.
Liu
,
M.
,
Wang
,
H.
,
Yoon
,
K.J.
and
Wang
,
L.
(
2024a
), “
Disentangled cross-modal fusion for event-guided image super-resolution
”,
IEEE Transactions on Artificial Intelligence
, Vol.
5
No.
10
.
Liu
,
S.
,
Li
,
J.
,
Zhao
,
G.
,
Zhang
,
Y.
,
Meng
,
X.
,
Yu
,
F.R.
,
Ji
,
X.
and
Li
,
M.
(
2025c
), “
EventGPT: event stream understanding with multimodal large language models
”,
Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR),
pp.
29139
-
29149
.
Liu
,
Y.
,
Deng
,
Y.
,
Chen
,
H.
and
Yang
,
Z.
(
2024b
), “
Video frame interpolation via direct synthesis with the event-based reference
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
8477
-
8487
.
Liu
,
Y.
,
Wei
,
L.
,
Guo
,
Y.
and
Yu
,
L.
(
2025a
), “
All-in-focus imaging from events with occlusions
”,
IEEE Transactions on Multimedia
, Vol.
27
.
Lou
,
H.
,
Teng
,
M.
,
Yang
,
Y.
and
Shi
,
B.
(
2023
), “
All-in-focus imaging from event focal stack
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
17366
-
17375
.
Low
,
W.F.
and
Lee
,
G.H.
(
2023
), “
Robust e-nerf: Nerf from sparse and noisy events under non-uniform motion
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
18335
-
18346
.
Low
,
W.F.
and
Lee
,
G.H.
(
2024
), “
Deblur e-NeRF: NeRF from motion-blurred events under high-speed or low-light conditions
”,
European Conference on Computer Vision
,
Springer
, pp.
192
-
209
.
Lu
,
Y.
,
Liang
,
G.
and
Wang
,
L.
(
2023a
), “
Self-supervised learning of event-guided video frame interpolation for rolling shutter frames
”,
IEEE Transactions on Visualization and Computer Graphics
, Vol.
31
No.
10
,
available at:
Link to Self-supervised learning of event-guided video frame interpolation for rolling shutter framesLink to the cited article.
Lu
,
Y.
,
Liang
,
G.
,
Wang
,
Y.
,
Wang
,
L.
and
Xiong
,
H.
(
2023b
), “
UniINR: event-guided unified rolling shutter correction, deblurring, and interpolation
”,
European Conference on Computer Vision
,
available at:
Link to UniINR: event-guided unified rolling shutter correction, deblurring, and interpolationLink to the cited article.
Lu
,
Y.
,
Qian
,
Y.
,
Rao
,
Z.
,
Xiao
,
J.
,
Chen
,
L.
and
Xiong
,
H.
(
2025
), “
RGBEvent ISP: the data set and benchmark
”,
arXiv preprint
.
Lu
,
Y.
,
Wang
,
Z.
,
Liu
,
M.
,
Wang
,
H.
and
Wang
,
L.
(
2023c
), “
Learning spatial-temporal implicit neural representations for event-guided video super-resolution
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
1557
-
1567
.
Ma
,
Q.
,
Paudel
,
D.P.
,
Chhatkuli
,
A.
and
Van Gool
,
L.
(
2023
), “
Deformable neural radiance fields using RGB and event cameras
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
3590
-
3600
.
Ma
,
Y.
,
Guo
,
S.
,
Chen
,
Y.
,
Xue
,
T.
and
Gu
,
J.
(
2024
), “
TimeLens-XL: real-time event-based video frame interpolation with large motion
”,
European Conference on Computer Vision
,
available at:
Link to TimeLens-XL: real-time event-based video frame interpolation with large motionLink to the cited article.
Maqueda
,
A.I.
,
Loquercio
,
A.
,
Gallego
,
G.
,
García
,
N.
and
Scaramuzza
,
D.
(
2018
), “
Event-based vision meets deep learning on steering prediction for self-driving cars
”,
Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
, pp.
5419
-
5427
.
Messikommer
,
N.
,
Gehrig
,
D.
,
Loquercio
,
A.
and
Scaramuzza
,
D.
(
2020
), “
Event-based asynchronous sparse convolutional networks
”,
European Conference on Computer Vision
,
Springer
, pp.
415
-
431
.
Messikommer
,
N.
,
Georgoulis
,
S.
,
Gehrig
,
D.
,
Tulyakov
,
S.
,
Erbach
,
J.
,
Bochicchio
,
A.
,
Li
,
Y.
and
Scaramuzza
,
D.
(
2022
), “
Multi-bracket high dynamic range imaging with event cameras
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
547
-
557
.
Mostafavi
,
M.
,
Wang
,
L.
,
Ho
,
Y.S.
and
Yoon
,
K.J.
(
2018
), “
Event-based high dynamic range image and very high frame rate video generation using conditional generative adversarial networks
”,
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
10073
-
10082
.
Mueggler
,
E.
,
Rebecq
,
H.
,
Gallego
,
G.
,
Delbruck
,
T.
and
Scaramuzza
,
D.
(
2017
), “
The event-camera data set and simulator: event-based data for pose estimation, visual odometry, and SLAM
”,
The International Journal of Robotics Research
, Vol.
36
No.
2
, pp.
142
-
149
.
Munda
,
G.
,
Reinbacher
,
C.
and
Pock
,
T.
(
2018
), “
Real-time intensity image reconstruction for event cameras using manifold regularization
”,
International Journal of Computer Vision
, Vol.
126
No.
12
, pp.
1381
-
1393
.
Musunuri
,
S.H.
,
Choudhury
,
A.
,
Su
,
G.M.
and
Park
,
S.
(
2024
), “
Event-based video frame extrapolation with recurrent neural networks
”,
2024 58th Asilomar Conference on Signals, Systems, and Computers
,
IEEE
, pp.
480
-
484
.
Nottebaum
,
M.
,
Roth
,
S.
and
Schaub-Meyer
,
S.
(
2022
), “
Efficient feature extraction for high-resolution video frame interpolation
”,
British Machine Vision Conference
,
available at:
Link to Efficient feature extraction for high-resolution video frame interpolationLink to the cited article.
Orchard
,
G.
,
Meyer
,
C.
,
Etienne-Cummings
,
R.
,
Posch
,
C.
,
Thakor
,
N.
and
Benosman
,
R.
(
2015
), “
HFirst: a temporal approach to object recognition
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
37
No.
10
, pp.
2028
-
2040
.
Palinauskas
,
G.
,
Amaya
,
C.
,
Eames
,
E.
,
Neumeier
,
M.
and
Von Arnim
,
A.
(
2023
), “
Generating event-based data sets for robotic applications using mujocoesim
”,
Proceedings of the 2023 International Conference on Neuromorphic Systems
, pp.
1
-
7
.
Pan
,
L.
,
Hartley
,
R.I.
,
Scheerlinck
,
C.
,
Liu
,
M.
,
Yu
,
X.
and
Dai
,
Y.
(
2019a
), “
High frame rate video reconstruction based on an event camera
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
44
No.
5
, pp.
2519
-
2533
.
Pan
,
L.
,
Scheerlinck
,
C.
,
Yu
,
X.
,
Hartley
,
R.
,
Liu
,
M.
and
Dai
,
Y.
(
2019b
), “
Bringing a blurry frame alive at high frame-rate with an event camera
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
6820
-
6829
.
Paredes-Vallés
,
F.
and
de Croon
,
G.C.
(
2020
), “
Back to event basics: self-supervised learning of image reconstruction for event cameras via photometric constancy
”,
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
3445
-
3454
,
available at:
Link to Back to event basics: self-supervised learning of image reconstruction for event cameras via photometric constancyLink to the cited article.
Park
,
J.
,
Moon
,
G.
,
Xu
,
W.
,
Kaseman
,
E.
,
Shiratori
,
T.
and
Lee
,
K.M.
(
2024
), “
3D hand sequence recovery from real blurry images and event stream
”,
European Conference on Computer Vision
,
Springer
, pp.
343
-
359
.
Peng
,
Y.
,
Zhang
,
Y.
,
Xiong
,
Z.
,
Sun
,
X.
and
Wu
,
F.
(
2023
), “
Get: group event transformer for event-based vision
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
6038
-
6048
.
Qi
,
Y.
,
Li
,
J.
,
Zhao
,
Y.
,
Zhang
,
Y.
and
Zhu
,
L.
(
2024a
), “
E3 NeRF: efficient event-enhanced neural radiance fields from blurry images
”,
arXiv preprint
.
Qi
,
Y.
,
Zhu
,
L.
,
Zhang
,
Y.
and
Li
,
J.
(
2023
), “
E2nerf: event enhanced neural radiance fields from blurry images
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
13254
-
13264
.
Qi
,
Y.
,
Zhu
,
L.
,
Zhao
,
Y.
,
Bao
,
N.
and
Li
,
J.
(
2024b
), “
Deblurring neural radiance fields with event-driven bundle adjustment
”,
Proceedings of the 32nd ACM International Conference on Multimedia
, pp.
9262
-
9270
.
Rebecq
,
H.
,
Gallego
,
G.
,
Mueggler
,
E.
and
Scaramuzza
,
D.
(
2018
), “
EMVS: event-based multi-view stereo – 3D reconstruction with an event camera in real-time
”,
International Journal of Computer Vision
, Vol.
126
No.
12
, pp.
1394
-
1414
.
Rebecq
,
H.
,
Ranftl
,
R.
,
Koltun
,
V.
and
Scaramuzza
,
D.
(
2019
), “
Eventsto-video: bringing modern computer vision to event cameras
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
3857
-
3866
.
Ruan
,
C.
,
Zhao
,
C.
,
Liang
,
C.
,
Luo
,
X.
,
Xu
,
J.
and
Chen
,
X.
(
2024
), “
Distill drops into data: event-based rain-background decomposition network
”,
Proceedings of the 30th Annual International Conference on Mobile Computing and Networking
, pp.
2072
-
2077
.
Rudnev
,
V.
,
Elgharib
,
M.
,
Theobalt
,
C.
and
Golyanik
,
V.
(
2023
), “
Eventnerf: neural radiance fields from a single color event camera
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
4992
-
5002
.
Rudnev
,
V.
,
Fox
,
G.
,
Elgharib
,
M.
,
Theobalt
,
C.
and
Golyanik
,
V.
(
2024
), “
Dynamic EventNeRF: reconstructing general dynamic scenes from multi-view event cameras
”,
arXiv preprint
.
Santambrogio
,
R.
,
Cannici
,
M.
and
Matteucci
,
M.
(
2024
), “
FARSE-CNN: fully asynchronous, recurrent and sparse event-based CNN
”,
European Conference on Computer Vision
,
Springer
, pp.
1
-
18
.
Scheerlinck
,
C.
,
Barnes
,
N.
and
Mahony
,
R.E.
(
2018
), “
Continuous-time intensity estimation using event cameras
”, ArXiv, ,
available at:
Link to Continuous-time intensity estimation using event camerasLink to the cited article.
Scheerlinck
,
C.
,
Rebecq
,
H.
,
Stoffregen
,
T.
,
Barnes
,
N.
,
Mahony
,
R.
and
Scaramuzza
,
D.
(
2019
), “
CED: color event camera data set
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops
.
Shariff
,
W.
,
Dilmaghani
,
M.S.
,
Kielty
,
P.
,
Moustafa
,
M.
,
Lemley
,
J.
and
Corcoran
,
P.
(
2024
), “
Event cameras in automotive sensing: a review
”,
IEEE Access
, Vol.
12
.
Shaw
,
R.
,
Catley-Chandar
,
S.
,
Leonardis
,
A.
and
Perez-Pellitero
,
E.
(
2022
), “
Hdr reconstruction from bracketed exposures and events
”,
arXiv preprint
.
Shi
,
C.
,
Liu
,
H.
,
Jin
,
J.
,
Li
,
W.
,
Li
,
Y.
,
Wei
,
B.
and
Zhang
,
Y.
(
2023
), “
IDOVFI: identifying dynamics via optical flow guidance for video frame interpolation with events
”, ArXiv, .
Shi
,
C.
,
Wei
,
B.
,
Wang
,
X.
,
Liu
,
H.
,
Zhang
,
Y.
,
Li
,
W.
,
Song
,
N.
and
Jin
,
J.
(
2024
), “
Polarity-focused denoising for event cameras
”,
IEEE Transactions on Circuits and Systems for Video Technology
, Vol.
35
No.
5
.
Sitzmann
,
V.
,
Martel
,
J.
,
Bergman
,
A.
,
Lindell
,
D.
and
Wetzstein
,
G.
(
2020
), “
Implicit neural representations with periodic activation functions
”,
Advances in Neural Information Processing Systems
, Vol.
33
, pp.
7462
-
7473
.
Song
,
C.
,
Huang
,
Q.X.
and
Bajaj
,
C.L.
(
2022
), “
E-CIR: event-enhanced continuous intensity recovery
”,
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
7793
-
7802
,
available at:
Link to E-CIR: event-enhanced continuous intensity recoveryLink to the cited article.
Stoffregen
,
T.
,
Scheerlinck
,
C.
,
Scaramuzza
,
D.
,
Drummond
,
T.
,
Barnes
,
N.
,
Kleeman
,
L.
and
Mahony
,
R.
(
2020
), “,“
Reducing the sim-toreal gap for event cameras
”,
Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVII 16
,
Springer
, pp.
534
-
549
.
Sun
,
L.
,
Sakaridis
,
C.
,
Liang
,
J.
,
Jiang
,
Q.
,
Yang
,
K.
,
Sun
,
P.
,
Ye
,
Y.
,
Wang
,
K.
and
Gool
,
L.V.
(
2022b
), “
Event-based fusion for motion deblurring with cross-modal attention
”,
European Conference on Computer Vision
,
Springer
, pp.
412
-
428
.
Sun
,
L.
,
Sakaridis
,
C.
,
Liang
,
J.
,
Sun
,
P.
,
Cao
,
J.
,
Zhang
,
K.
,
Jiang
,
Q.
,
Wang
,
K.
and
Van Gool
,
L.
(
2023
), “
Event-based frame interpolation with ad-hoc deblurring
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
18043
-
18052
.
Sun
,
L.
,
Zhang
,
Y.
,
Cheng
,
K.
,
Cheng
,
J.
and
Lu
,
H.
(
2022a
), “
MeNet: a memory-based network with dual-branch for efficient event stream processing
”,
European Conference on Computer Vision
,
Springer
, pp.
214
-
234
.
Sun
,
Z.
,
Fu
,
X.
,
Huang
,
L.
,
Liu
,
A.
and
Zha
,
Z.J.
(
2024
), “
Motion aware event representation-driven image deblurring
”,
European Conference on Computer Vision
,
available at:
Link to Motion aware event representation-driven image deblurringLink to the cited article.
Tang
,
J.
,
Lai
,
J.H.
,
Yang
,
L.
and
Xie
,
X.
(
2025
), “
Spike-temporal latent representation for energy-efficient event-to-video reconstruction
”,
European Conference on Computer Vision
,
Springer
, pp.
163
-
179
.
Tavanaei
,
A.
,
Ghodrati
,
M.
,
Kheradpisheh
,
S.R.
,
Masquelier
,
T.
and
Maida
,
A.
(
2019
), “
Deep learning in spiking neural networks
”,
Neural Networks: The Official Journal of the International Neural Network Society
, Vol.
111
, pp.
47
-
63
.
Teng
,
M.
,
Lou
,
H.
,
Yang
,
Y.
,
Huang
,
T.
and
Shi
,
B.
(
2024
), “
Hybrid Allin-focus imaging from neuromorphic focal stack
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
46
No.
12
.
Teng
,
M.
,
Zhou
,
C.
,
Lou
,
H.
and
Shi
,
B.
(
2022
), “
NEST: neural event stack for event-based image enhancement
”,
European Conference on Computer Vision
,
Springer
, pp.
660
-
676
.
Tian
,
D.
,
Choudhury
,
A.
and
Husak
,
W.
(
2025
), “
Long-short exposure fusion with event data for low-light video enhancement
”,
2025 IEEE International Conference on Image Processing (ICIP)
,
IEEE
, pp.
1660
-
1665
.
Todorov
,
E.
,
Erez
,
T.
and
Tassa
,
Y.
(
2012
), “
MuJoCo: a physics engine for model-based control
”,
2012 IEEE/RSJ International Conference on Intelligent Robots and Systems
,
IEEE
, pp.
5026
-
5033
, doi: .
Tulyakov
,
S.
,
Bochicchio
,
A.
,
Gehrig
,
D.
,
Georgoulis
,
S.
,
Li
,
Y.
and
Scaramuzza
,
D.
(
2022
), “
Time lens++: event-based frame interpolation with parametric non-linear flow and multi-scale fusion
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
17755
-
17764
.
Tulyakov
,
S.
,
Gehrig
,
D.
,
Georgoulis
,
S.
,
Erbach
,
J.
,
Gehrig
,
M.
,
Li
,
Y.
and
Scaramuzza
,
D.
(
2021
), “
Time lens: event-based video frame interpolation
”,
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
, pp.
16155
-
16164
.
Vaswani
,
A.
,
Shazeer
,
N.
,
Parmar
,
N.
,
Uszkoreit
,
J.
,
Jones
,
L.
,
Gomez
,
A.N.
,
Kaiser
,
Ł.
and
Polosukhin
,
I.
(
2017
), “
Attention is all you need
”,
Advances in Neural Information Processing Systems
, p.
30
.
Wang
,
B.
,
He
,
J.
,
Yu
,
L.
,
Xia
,
G.S.
and
Yang
,
W.
(
2020
), “
Event enhanced high-quality image recovery
”,
Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIII 16
,
Springer
, pp.
155
-
171
.
Wang
,
J.
,
He
,
J.
,
Zhang
,
Z.
and
Xu
,
R.
(
2024c
), “
Physical priors augmented event-based 3d reconstruction
”,
2024 IEEE International Conference on Robotics and Automation (ICRA)
,
IEEE
, pp.
16810
-
16817
.
Wang
,
J.
,
Weng
,
W.
,
Zhang
,
Y.
and
Xiong
,
Z.
(
2023a
), “
Unsupervised video deraining with an event camera
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
10831
-
10840
.
Wang
,
L.
,
Ho
,
Y.S.
,
Yoon
,
K.J.
and
Mostafavi
,
M.
(
2019
), “
Event-based high dynamic range image and very high frame rate video generation using conditional generative adversarial networks
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
10081
-
10090
.
Wang
,
P.
,
He
,
J.
,
Yan
,
Q.
,
Zhu
,
Y.
,
Sun
,
J.
and
Zhang
,
Y.
(
2024e
), “
Diffevent: event residual diffusion for image deblurring
”,
ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
, pp.
3450
-
3454
,
available at:
Link to Diffevent: event residual diffusion for image deblurringLink to the cited article.
Wang
,
R.
,
Guo
,
Q.
,
Li
,
H.
and
Wan
,
R.
(
2024b
), “
Event Trojan: asynchronous event-based backdoor attacks
”,
European Conference on Computer Vision
,
Springer
, pp.
315
-
332
.
Wang
,
X.
,
Fu
,
H.
,
Wang
,
J.
,
Wang
,
X.
,
Zhang
,
H.
and
Ma
,
H.
(
2024d
), “
Exploring in extremely dark: low-light video enhancement with real events
”,
Proceedings of the 32nd ACM International Conference on Multimedia
, pp.
4805
-
4813
.
Wang
,
Z.
,
Hamann
,
F.
,
Chaney
,
K.
,
Jiang
,
W.
,
Gallego
,
G.
and
Daniilidis
,
K.
(
2023b
), “
Event-based continuous color video decompression from single frames
”, ArXiv, ,
available at:
Link to Event-based continuous color video decompression from single framesLink to the cited article.
Wang
,
Z.
,
Lu
,
Y.
and
Wang
,
L.
(
2024a
), “
Revisit event generation model: self-supervised learning of event-to-video reconstruction with implicit neural representations
”,
European Conference on Computer Vision
,
available at:
Link to Revisit event generation model: self-supervised learning of event-to-video reconstruction with implicit neural representationsLink to the cited article.
Weng
,
J.
,
Li
,
B.
and
Huang
,
K.
(
2024a
), “
Event-based image enhancement under high dynamic range scenarios
”,
Proceedings of the Asian Conference on Computer Vision
, pp.
2456
-
2470
.
Weng
,
W.
,
Zhang
,
Y.
and
Xiong
,
Z.
(
2021
), “
Event-based video reconstruction using transformer
”,
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
, pp.
2543
-
2552
,
available at:
Link to Event-based video reconstruction using transformerLink to the cited article.
Weng
,
W.
,
Zhang
,
Y.
and
Xiong
,
Z.
(
2022
), “
Boosting event stream super-resolution with a recurrent neural network
”,
European Conference on Computer Vision
,
Springer
, pp.
470
-
488
.
Weng
,
W.
,
Zhang
,
Y.
and
Xiong
,
Z.
(
2023
), “
Event-based blurry frame interpolation under blind exposure
”,
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
1588
-
1598
,
available at:
Link to Event-based blurry frame interpolation under blind exposureLink to the cited article.
Weng
,
Y.
,
Shen
,
Z.
,
Chen
,
R.
,
Wang
,
Q.
and
Wang
,
J.
(
2024b
), “
Eadeblurgs: event assisted 3d deblur reconstruction with Gaussian splatting
”,
arXiv preprint
.
Wu
,
S.
,
You
,
K.
,
He
,
W.
,
Yang
,
C.
,
Tian
,
Y.
,
Wang
,
Y.
,
Zhang
,
Z.
and
Liao
,
J.
(
2022
), “
Video interpolation by event-driven anisotropic adjustment of optical flow
”,
European Conference on Computer Vision
.
Wu
,
Y.
,
Tan
,
G.
,
Chen
,
J.
,
Zhai
,
W.
,
Cao
,
Y.
and
Zha
,
Z.J.
(
2024
), “
Eventbased asynchronous HDR imaging by temporal incident light modulation
”,
Optics Express
, Vol.
32
No.
11
, pp.
18527
-
18538
.
Xia
,
L.
,
Xiong
,
R.
,
Zhao
,
J.
,
Wang
,
L.
,
Zhu
,
S.
,
Fan
,
X.
and
Huang
,
T.
(
2025
), “
High spatio-temporal imaging reconstruction for hybrid spike-RGB cameras
”,
IEEE Transactions on Computational Imaging
, Vol.
11
, pp.
586
-
598
.
Xia
,
L.
,
Zhao
,
J.
,
Xiong
,
R.
and
Huang
,
T.
(
2023
), “
SVFI: spiking-based video frame interpolation for high-speed motion
”,
Proceedings of the AAAI Conference on Artificial Intelligence
, Vol.
37
No.
3
, pp.
2910
-
2918
.
Xiang
,
X.
,
Zhu
,
L.
,
Li
,
J.
,
Tian
,
Y.
and
Huang
,
T.
(
2022
), “
Temporal upsampling for asynchronous events
”,
2022 IEEE International Conference on Multimedia and Expo (ICME)
,
IEEE
, pp.
1
-
6
.
Xiao
,
Z.
and
Wang
,
X.
(
2025
), “
Event-based video super-resolution via state space models
”,
Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR),
pp.
12564
-
12574
.
Xiao
,
Z.
,
Kai
,
D.
,
Zhang
,
Y.
,
Sun
,
X.
and
Xiong
,
Z.
(
2024a
), “
Asymmetric event-guided video super-resolution
”,
Proceedings of the 32nd ACM International Conference on Multimedia
, pp.
2409
-
2418
.
Xiao
,
Z.
,
Kai
,
D.
,
Zhang
,
Y.
,
Zha
,
Z.J.
,
Sun
,
X.
and
Xiong
,
Z.
(
2024b
), “
Event-adapted video super-resolution
”,
European Conference on Computer Vision
,
Springer
, pp.
217
-
235
.
Xiaopeng
,
L.
,
Zhaoyuan
,
Z.
,
Cien
,
F.
,
Chen
,
Z.
,
Lei
,
D.
and
Lei
,
Y.
(
2024
), “
HDR imaging for dynamic scenes with events
”,
arXiv preprint
.
Xie
,
X.
,
Zhang
,
Q.
and
Zheng
,
W.S.
(
2025
), “
Diffusion-based event generation for high-quality image deblurring
”,
Proceedings of the Computer Vision and Pattern Recognition Conference
, pp.
2194
-
2203
.
Yang
,
Y.
,
Han
,
J.
,
Liang
,
J.
,
Sato
,
I.
and
Shi
,
B.
(
2023b
), “
Learning event guided high dynamic range video reconstruction
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
13924
-
13934
.
Yang
,
Y.
,
Liang
,
J.
,
Yu
,
B.
,
Chen
,
Y.
,
Ren
,
J.S.
and
Shi
,
B.
(
2024
), “
Latency correction for event-guided deblurring and frame interpolation
”,
2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
24977
-
24986
,
available at:
Link to Latency correction for event-guided deblurring and frame interpolationLink to the cited article.
Yang
,
Y.
,
Pan
,
L.
and
Liu
,
L.
(
2023a
), “
Event camera data pre-training
”,
Proceedings of the IEEE/CVF International Conference on Computer Vision
, pp.
10699
-
10709
.
Yao
,
Y.
,
Zhao
,
X.
and
Gu
,
B.
(
2024
), “
Exploring vulnerabilities in spiking neural networks: direct adversarial attacks on raw event data
”,
European Conference on Computer Vision
,
Springer
, pp.
412
-
428
.
Yu
,
B.
,
Ren
,
J.
,
Han
,
J.
,
Wang
,
F.
,
Liang
,
J.
and
Shi
,
B.
(
2024a
), “
EventPS: real-time photometric stereo using an event camera
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
9602
-
9611
.
Yu
,
L.
,
Wang
,
B.
,
Zhang
,
X.
,
Zhang
,
H.
,
Yang
,
W.
,
Liu
,
J.
and
Xia
,
G.S.
(
2023
), “
Learning to super-resolve blurry images with events
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
45
no No.
8
, pp.
10027
-
10043
.
Yu
,
L.
,
Zhang
,
X.
,
Liao
,
W.
,
Yang
,
W.
and
Xia
,
G.S.
(
2022
), “
Learning to see through with events
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
45
no No.
7
, pp.
8660
-
8678
.
Yu
,
W.
,
Feng
,
C.
,
Tang
,
J.
,
Yang
,
J.
,
Tang
,
Z.
,
Jia
,
X.
,
Yang
,
Y.
,
Yuan
,
L.
and
Tian
,
Y.
(
2024b
), “
EvaGaussians: event stream assisted Gaussian splatting from blurry images
”,
arXiv preprint
.
Yunfan
,
L.
,
Xu
,
X.
,
Hao
,
L.
,
Qian
,
Y.
,
Yang
,
B.
,
Li
,
J.
,
Cai
,
Q.
,
Guo
,
W.
and
Xiong
,
H.
(
2025
), “
SEE: see everything every time-broader light range image enhancement via events
”.
Yura
,
T.
,
Mirzaei
,
A.
and
Gilitschenski
,
I.
(
2025
), “
EventSplat: 3D Gaussian splatting from moving event cameras for real-time rendering
”,
Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR),
pp.
26876
-
26886
.
Zhang
,
B.
,
Han
,
Y.
,
Suo
,
J.
and
Dai
,
Q.
(
2024a
), “
An event-oriented diffusion refinement method for sparse events completion
”,
Scientific Reports
, Vol.
14
No.
1
, p.
6802
.
Zhang
,
C.
,
Lin
,
M.
,
Zhang
,
X.
,
Jiang
,
C.
and
Yu
,
L.
(
2024b
), “
Super-resolving blurry images with events
”,
arXiv preprint
.
Zhang
,
C.
,
Zhang
,
X.
,
Lin
,
M.
,
Li
,
C.
,
He
,
C.
,
Yang
,
W.
,
Xia
,
G.S.
and
Yu
,
L.
(
2024d
), “
CrossZoom: simultaneous motion deblurring and event super-resolving
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
46
No.
12
.
Zhang
,
C.
,
Zhang
,
X.
,
Lin
,
M.
,
Li
,
C.
,
He
,
C.
,
Yang
,
W.
,
Xia
,
G.S.
and
Yu
,
L.
(
2023f
), “
CrossZoom: simultaneous motion deblurring and event super-resolving
”,
IEEE Transactions on Pattern Analysis and Machine Intelligence
, Vol.
46
No.
12
, pp.
8209
-
8227
,
available at:
Link to CrossZoom: simultaneous motion deblurring and event super-resolvingLink to the cited article.
Zhang
,
G.
,
Zhu
,
Y.
,
Wang
,
H.
,
Chen
,
Y.
,
Wu
,
G.
and
Wang
,
L.
(
2023e
), “
Extracting motion and appearance via inter-frame attention for efficient video frame interpolation
”,
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
5682
-
5692
.
Zhang
,
J.
,
Chen
,
S.
,
Zheng
,
Y.
,
Yu
,
Z.
and
Huang
,
T.
(
2023a
), “
Unveiling the potential of spike streams for foreground occlusion removal from densely continuous views
”,
arXiv preprint
.
Zhang
,
J.
,
Chen
,
S.
,
Zheng
,
Y.
,
Yu
,
Z.
and
Huang
,
T.
(
2024b
), “
Transient glimpses: unveiling occluded backgrounds through the spike camera
”,
Proceedings of the AAAI Conference on Artificial Intelligence
, Vol.
38
No.
1
, pp.
637
-
645
.
Zhang
,
J.
,
Chen
,
S.
,
Zheng
,
Y.
,
Yu
,
Z.
and
Huang
,
T.
(
2024c
), “
Spikeguided motion deblurring with unknown modal spatiotemporal alignment
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
25047
-
25057
.
Zhang
,
P.
,
Liu
,
H.
,
Ge
,
Z.
,
Wang
,
C.
and
Lam
,
E.Y.
(
2024d
), “
Neuromorphic imaging with joint image deblurring and event denoising
”,
IEEE Transactions on Image Processing
, Vol.
33
.
Zhang
,
S.
,
Zhang
,
Y.
,
Jiang
,
Z.
,
Zou
,
D.
,
Ren
,
J.
and
Zhou
,
B.
(
2020
), “
Learning to see in the dark with events
”,
Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVIII 16
,
Springer
, pp.
666
-
682
.
Zhang
,
X.
and
Yu
,
L.
(
2022
), “
Unifying motion deblurring and frame interpolation with events
”,
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
17744
-
17753
,
available at:
Link to Unifying motion deblurring and frame interpolation with eventsLink to the cited article.
Zhang
,
X.
,
Liao
,
W.
,
Yu
,
L.
,
Yang
,
W.
and
Xia
,
G.S.
(
2021
), “
Eventbased synthetic aperture imaging with a hybrid network
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
14235
-
14244
.
Zhang
,
X.
,
Yu
,
L.
,
Yang
,
W.
,
Liu
,
J.Z.
and
Xia
,
G.S.
(
2023d
), “
Generalizing event-based motion deblurring in real-world scenarios
”,
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
, pp.
10700
-
10710
,
available at:
Link to Generalizing event-based motion deblurring in real-world scenariosLink to the cited article.
Zhang
,
Z.
,
Cui
,
S.
,
Chai
,
K.
,
Yu
,
H.
,
Dasgupta
,
S.
,
Mahbub
,
U.
and
Rahman
,
T.
(
2024
g), “
V2ce: video to continuous events simulator
”,
2024 IEEE International Conference on Robotics and Automation (ICRA)
,
IEEE
, pp.
12455
-
12461
.
Zhang
,
Z.
,
Ma
,
Y.
,
Chen
,
Y.
,
Zhang
,
F.
,
Gu
,
J.
,
Xue
,
T.
and
Guo
,
S.
(
2024f
), “
From sim-to-real: toward general event-based low light frame interpolation with per-scene optimization
”,
SIGGRAPH Asia 2024 Conference Papers
, pp.
1
-
10
.
Zheng
,
X.
,
Liu
,
Y.
,
Lu
,
Y.
,
Hua
,
T.
,
Pan
,
T.
,
Zhang
,
W.
,
Tao
,
D.
and
Wang
,
L.
(
2023a
), “
Deep learning for event-based vision: a comprehensive survey and benchmarks
”,
arXiv preprint
.
Zheng
,
Y.
,
Zhang
,
J.
,
Zhao
,
R.
,
Ding
,
J.
,
Chen
,
S.
,
Xiong
,
R.
,
Yu
,
Z.
and
Huang
,
T.
(
2023b
), “
SpikeCV: open a continuous computer vision era
”,
arXiv preprint
.
Zhou
,
Y.
,
Gallego
,
G.
,
Rebecq
,
H.
,
Kneip
,
L.
,
Li
,
H.
and
Scaramuzza
,
D.
(
2018
), “
Semi-dense 3D reconstruction with a stereo event camera
”,
Proceedings of the European Conference on Computer Vision (ECCV),
pp.
235
-
251
.
Zhu
,
A.Z.
,
Thakur
,
D.
,
Özaslan
,
T.
,
Pfrommer
,
B.
,
Kumar
,
V.
and
Daniilidis
,
K.
(
2018
), “
The multivehicle stereo event camera data set: an event camera data set for 3D perception
”,
IEEE Robotics and Automation Letters
, Vol.
3
No.
3
, pp.
2032
-
2039
.
Zhu
,
L.
,
Wang
,
X.
,
Chang
,
Y.
,
Li
,
J.
,
Huang
,
T.
and
Tian
,
Y.
(
2022
), “
Event-based video reconstruction via potential-assisted spiking neural network
”,
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
3584
-
3594
,
available at:
Link to Event-based video reconstruction via potential-assisted spiking neural networkLink to the cited article.
Zhu
,
L.
,
Zheng
,
Y.
,
Zhang
,
Y.
,
Wang
,
X.
,
Wang
,
L.
and
Huang
,
H.
(
2025
), “
Temporal residual guided diffusion framework for event-driven video reconstruction
”,
European Conference on Computer Vision
,
Springer
, pp.
411
-
427
.
Ziegler
,
A.
,
Teigland
,
D.
,
Tebbe
,
J.
,
Gossard
,
T.
and
Zell
,
A.
(
2023
), “
Real time event simulation with frame-based cameras
”,
2023 IEEE International Conference on Robotics and Automation (ICRA)
,
IEEE
, pp.
11669
-
11675
.
Zihao Zhu
,
A.
,
Yuan
,
L.
,
Chaney
,
K.
and
Daniilidis
,
K.
(
2018
), “
Unsupervised event-based optical flow using motion compensation
”,
Proceedings of the European Conference on Computer Vision (ECCV) Workshops
.
Zou
,
Y.
,
Zheng
,
Y.
,
Takatani
,
T.
and
Fu
,
Y.
(
2021
), “
Learning to reconstruct high speed and high dynamic range videos from events
”,
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
, pp.
2024
-
2033
,
available at:
Link to Learning to reconstruct high speed and high dynamic range videos from eventsLink to the cited article.
Zubic
,
N.
,
Gehrig
,
M.
and
Scaramuzza
,
D.
(
2024
), “
State space models for event cameras
”,
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
, pp.
5819
-
5828
.
Liu
,
H.C.
,
Zhang
,
F.L.
,
Marshall
,
D.
,
Shi
,
L.
and
Hu
,
S.M.
(
2017
), “
High-speed video generation with an event camera
”,
The Visual Computer
, Vol.
33
Nos
6-8
, pp.
749
-
759
.
Zhang
,
P.
,
Liu
,
H.
,
Ge
,
Z.
,
Wang
,
C.
and
Lam
,
E.Y.
(
2023c
), “
Neuromorphic imaging with joint image deblurring and event denoising
”,
IEEE Transactions on Image Processing: a Publication of the IEEE Signal Processing Society
, Vol.
33
, pp.
2318
-
2333
,
available at:
Link to Neuromorphic imaging with joint image deblurring and event denoisingLink to the cited article.
Zhang
,
X.
,
Huang
,
H.
,
Jia
,
X.
,
Wang
,
D.
and
Lu
,
H.
(
2023b
), “
Neural image re-exposure
”, ArXiv, ,
available at:
Link to Neural image re-exposureLink to the cited article.
Zhang
,
Y.
,
Wang
,
J.
,
Weng
,
W.
,
Sun
,
X.
and
Xiong
,
Z.
(
2023c
), “
EGVD: event-guided video deraining
”,
arXiv preprint
.
Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at Link to the terms of the CC BY 4.0 licenceLink to the terms of the CC BY 4.0 licence

or Create an Account

Close subscription notice
Close access options