Table 1

Comparative analysis of vision-based HIP models

ModelsFeatures
Two-stream-based CNNs (Li et al., 2020)Human action is determined by calculating the L2 distance of the positions of the human joints between frames (CNNs)
CLSTM (Sarabu and Santra, 2021)Present a two-stream network with two CNNs and Convolution Long-Short Term Memory (CLSTM) (CNN + LSTM) a powerful feature extractor in human action recognition in videos
CLSTDN (Saif et al., 2023)This research proposes Convolutional Long Short-Term Deep Network (CLSTDN) consists of Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) for Recognition of human intention
Transformer and Bi-LSTM (Zhang et al., 2024)extract features from human motion trajectories by analyzing changes in human joint distances with a Transformer and a Bi-LSTM, respectively
GRU-CNN (Du et al., 2025)proposed multi-channel parallel GRU-CNN neural network combines the temporal analysis capabilities of GRU with the spatial feature extraction strengths of CNN through a weight allocation strategy for trajectory prediction and intention recognition
Source(s): Authors’ own work

or Create an Account

Close Modal
Close Modal