Geotechnical engineering is widely recognized as one of the most intricate fields of engineering due to its involvement with earth materials such as soil, rock and intermediate geo-materials, e.g. coal, which are associated with numerous sources of uncertainty. Consequently, physical studies in this field often rely on simplifications and empirical assumptions. However, the increasing availability of data in geotechnical projects and the potential use of deep learning (DL) algorithms to address research challenges in geotechnics have rendered DL a pivotal research subject in geotechnical engineering. Compared to traditional machine learning (ML) approaches, DL algorithms have greater capacity for automated feature extraction and for learning from large and complex data sets. Therefore, DL algorithms have gained broad application across diverse geotechnical domains. This state-of-the-art review is presented in two parts. Part 1 discusses the most commonly used DL algorithms based on the available literature. Part 2 then highlights the specific application of these algorithms in geotechnical projects. In the first part of the paper, the frequency of use for each algorithm is listed based on the number of articles in which they are used. Subsequently, the most important algorithms are examined in detail. This paper also aims to address the computational tools and frameworks used for implementing DL models, as well as the efficiency and scalability of DL algorithms. In addition, a comprehensive summary is provided, which includes published literature, relevant reference materials, adopted DL algorithms and relevant geotechnical topics. The paper concludes by evaluating the challenges and prospects that lie ahead for the future development of DL in geotechnical engineering.
This review adopts a systematic literature analysis of DL algorithms relevant to geotechnical engineering. It categorizes DL methods based on usage frequency and evaluates their structures, implementation frameworks (e.g. TensorFlow, Keras, PyTorch) and learning strategies, such as transfer learning. The paper emphasizes the role of data quality and preprocessing and compiles open-access data sets for researchers. It serves as Part 1 of a two-part series, focusing on DL techniques rather than specific applications.
Convolutional neural networks, long short-term memory networks and generative adversarial networks are among the most frequently applied DL models in geotechnical research. These models have demonstrated superior feature extraction, accuracy and adaptability to geotechnical data sets. DL frameworks like TensorFlow and PyTorch are widely used due to their scalability and open-source nature. Transfer learning emerges as a powerful technique in data-scarce geotechnical domains.
This is one of the first structured reviews exclusively dedicated to DL algorithms in geotechnical engineering. Unlike previous artificial intelligence surveys, it focuses on algorithmic foundations, tools and data requirements. The inclusion of transfer learning and publicly available data sets adds significant practical value. It serves both as an educational entry point and a foundation for further research into real-world DL applications in geotechnics.
1. Introduction
In recent decades, artificial intelligence (AI) has shown a notable influence across various disciplines of science (Du and Che, 2023). When examining the literature on a specific research topic, e.g. geotechnical engineering, it is evident that the adoption of these applications is made possible through the use of computational tools, specifically machine learning (ML) and deep learning (DL) methods (Phoon and Zhang, 2023; Cheng and Ziotopoulou, 2023). In general, the relationship between AI, ML and DL can be viewed from two different perspectives. First, as demonstrated in Figure 1(a), DL is a subset of ML, and ML is a subset of AI (Taye, 2023, Liu et al., 2023b; Nyarko et al., 2023). While this classification is not as absolute as natural laws, it is widely accepted. To further examine this relationship, a specific clarification about the meaning of these terms is needed.
Schematic representation of the relation between AI, ML and DL
Schematic representation of the relation between AI, ML and DL
AI is a broad term that describes the development of intelligent systems capable of performing tasks that often require human intelligence (Haenlein and Kaplan, 2019). It includes a wide range of technologies and approaches that attempt to mimic human cognitive abilities (Mehak and Ashima, 2023). ML is a subfield of AI that focuses on developing algorithms and models that enable computers to learn from data and make predictions or decisions without being explicitly programmed to do so (Andriole, 2019). ML algorithms learn patterns and relationships from data, allowing them to enhance performance over time (Sarker, 2021). DL is a subset of ML in which artificial neural networks with multiple layers are used to process and learn from large amounts of data (Choudhary et al., 2022). DL algorithms are intended to automatically extract complex features from data, allowing them to solve more difficult tasks (Najafabadi et al., 2015). Another distinction between ML and DL lies in the feature selection capability. In traditional ML systems, feature selection requires manual engagement from the user as a preliminary step, acting as an important preprocessing requirement before engaging with the main mechanisms of ML. In contrast, within the area of DL, this procedure is smoothly integrated within the operations of the deep neural network itself (Taye, 2023; Munappy et al., 2022). Specifically, feature selection and subsequent regression or classification tasks are fundamental components inherent in the core operations of DL algorithms [Figure 1(b)]. This inherent characteristic renders deep neural networks particularly well-suited for tasks involving complex, high-dimensional data, including image and video processing (Alzubaidi et al., 2021; Choudhary et al., 2022).
Geotechnical engineering, as one of the most important fields of civil engineering, concerns the understanding and management of the behavior of soil and rock to ensure the stability and durability of infrastructure projects (Yaghoubi and Yaghoubi, 2024; Robbins et al., 2021; Witold Bogusz, 2017). This field covers a wide range of tasks, from evaluation of soil parameters and mitigating the natural hazards to designing geo-structures. Historically, geotechnical engineers dealt with these challenges using empirical formulas, hand computations and basic models. However, these classic methodologies are frequently insufficient for comprehending the intricate interaction of geological, hydrological and environmental factors that shape soil behavior (Witold Bogusz, 2017; Bogusz and Godlewski, 2019). These limitations increase the demand for more innovative tools to tackle the increasing complexity of geotechnical analysis. In recent years, the advent of DL methods has influenced geotechnical engineering and changed the landscape of geotechnical engineering (Baghbani et al., 2022; Basu et al., 2015). DL uses the power of deep neural networks to extract intricate patterns and characteristics from large data sets. This shift has improved modeling and predicting in this area, and further accelerated decision-making processes by rapidly analyzing large amounts of data (Zhang et al., 2022; Zhang and Phoon, 2022).
Because many DL techniques originated in fields outside geotechnical engineering, the geotechnical community has not yet achieved a comprehensive understanding of their structure, usage and applications (Zhang et al., 2021b; Demertzis et al., 2023). In contrast, the pace of advancement in AI, and especially in DL, is rapid. As such, the results of these advances are introduced to the field of geotechnical engineering with a time lag, compounding the challenge for people involved in the domain to understand them (Onyelowe et al., 2023; Jaksa and Liu, 2021). It is thus imperative to conduct a comprehensive review of the DL techniques that have gained traction in geotechnical engineering to help readers see the main categories of DL methods and understand where each of them is typically applied (Ebid, 2020).
This paper first conducts a bibliometric analysis to practically find the most common DL algorithms in geotechnical engineering. It then introduces the main computational frameworks, namely, TensorFlow, Keras and PyTorch, that make it possible to build and train these models in practice. The core DL algorithms themselves are then explained in detail, along with their variations and practical relevance. As high-quality data is essential for any DL application, the paper also reviews the major data sources available to geotechnical engineers and the preprocessing steps required to prepare them. The paper is finally concluded by a summary and discussion on the main content.
2. Bibliometric analysis
In this section, a bibliometric analysis is conducted to quantitatively analyze the trends in published research on DL in geotechnical engineering. To this end, the Scopus database is used as a source of information. The search process is conducted using the keywords “Deep Learning” AND (“geotechnic*” OR “soil” OR “rock”) to extract papers related to the application of DL in geotechnical issues. In total, 6,470 articles are obtained from the search, and a statistical analysis of these articles provides interesting information. First, the publication trend of these papers and the total number of citations to the articles are shown in Figure 2 for the period from 2012 to 2025.
The horizontal axis is labelled Year and lists 2012 through 2025. The left vertical axis is labelled Number of Documents and ranges from 0 to 2000 with tick marks at 500 intervals. The right vertical axis is labelled Number of Citations and ranges from 0 to 3.5 times 10 to the power of 4, with the multiplier shown as times 10 to the power of 4. The legend identifies blue bars as Publications and a black line with circular markers as Citations. Publications remain at zero from 2012 to 2016, begin with a very small value in 2017, increase slightly in 2018 and 2019, then rise steadily through 2020, 2021, 2022, 2023, 2024, and reach the highest value in 2025 at about 1700 documents. Citations remain close to zero from 2012 to 2017, increase gradually in 2018, 2019, and 2020, then rise sharply through 2021, 2022, 2023, and 2024, reaching the highest value in 2025 at about 3.1 times 10 to the power of 4 citations.Trends in paper publications (blue bars) and citations (black line) in the field of deep learning in geotechnical engineering from 2012 to 2025
The horizontal axis is labelled Year and lists 2012 through 2025. The left vertical axis is labelled Number of Documents and ranges from 0 to 2000 with tick marks at 500 intervals. The right vertical axis is labelled Number of Citations and ranges from 0 to 3.5 times 10 to the power of 4, with the multiplier shown as times 10 to the power of 4. The legend identifies blue bars as Publications and a black line with circular markers as Citations. Publications remain at zero from 2012 to 2016, begin with a very small value in 2017, increase slightly in 2018 and 2019, then rise steadily through 2020, 2021, 2022, 2023, 2024, and reach the highest value in 2025 at about 1700 documents. Citations remain close to zero from 2012 to 2017, increase gradually in 2018, 2019, and 2020, then rise sharply through 2021, 2022, 2023, and 2024, reaching the highest value in 2025 at about 3.1 times 10 to the power of 4 citations.Trends in paper publications (blue bars) and citations (black line) in the field of deep learning in geotechnical engineering from 2012 to 2025
An analysis of the publication and citation trends shows that the application of AI in geotechnical engineering has grown significantly over the past decade. From 2012 to 2016, the number of published papers was very low, ranging from 1 to 9 articles per year; this period can be interpreted as the “initial stage or formation of research activity” in this field. Between 2017 and 2019, a relatively mild trend in increasing scientific production is observed (32, 78 and 158 articles per year, respectively), but from 2020 onwards, growth became much more intense, such that the number of papers increased from 338 in 2020 to 814 in 2022 and finally to 1,700 in 2025. In other words, in about five years, the volume of publications in this field has more than fivefold, indicating that this field has emerged as a rapidly developing research area. The trend in citations has also been in line with the increase in the number of papers, albeit with a time lag (citation lag). Until 2018, the number of citations was insignificant, but from 2019 to 2023, a significant increase is observed, reaching 30,683 in 2025. The citation-to-article ratio has also increased from about 2 to 4 in the early years to more than 18 citations per article in 2025, indicating a substantial increase in both publication output and citation activity.
After extracting the keywords of all papers obtained from Scopus, the keywords are separated, cleaned and aggregated into a unified data set. Subsequently, these keywords are analyzed to identify the most frequently used DL algorithms in geotechnical research. In this process, similar or synonymous algorithm names, e.g. CNN, Convolutional Neural Network and Conv, were consolidated into a single category to obtain the true final count for each algorithm. The results are reported in Table 1.
Frequency of commonly identified DL models and approaches in geotechnical publications
| No. | Algorithm | Keyword counts |
|---|---|---|
| 1 | Convolutional neural network (CNN) | 696 |
| 2 | Long short-term memory (LSTM) | 246 |
| 3 | Recurrent neural network (RNN) | 76 |
| 4 | U-Net / UNet | 84 |
| 5 | Transformer / attention | 157 |
| 6 | Generative adversarial network (GAN) | 38 |
| 7 | YOLO | 64 |
| 8 | ResNet | 53 |
| 9 | Deep neural network (DNN) | 86 |
| 10 | Deep reinforcement learning | 36 |
| No. | Algorithm | Keyword counts |
|---|---|---|
| 1 | Convolutional neural network ( | 696 |
| 2 | Long short-term memory ( | 246 |
| 3 | Recurrent neural network ( | 76 |
| 4 | U-Net / UNet | 84 |
| 5 | Transformer / attention | 157 |
| 6 | Generative adversarial network ( | 38 |
| 7 | 64 | |
| 8 | ResNet | 53 |
| 9 | Deep neural network ( | 86 |
| 10 | Deep reinforcement learning | 36 |
As can be seen from the keyword frequencies presented in the table, CNNs are by far the most widely used DL algorithm in this field. This is due to their ability to extract complex features from visual and spatial data, such as images, 3D earth models and remote sensing data. In addition to CNNs, LSTM networks and models based on the attention mechanism (i.e. Transformer/Attention) also make a significant contribution to geotechnical studies, which reflects the tendency of researchers to model time series and sequential data, such as groundwater level changes, soil stresses and displacements and other time-varying parameters. RNNs and DNNs also play an important role in analyzing nonlinear and complex data, although their use is less than that of CNN and LSTM. These algorithms provide flexible tools for modeling nonlinear relationships and predicting geotechnical responses.
3. Computational tools and frameworks in DL
Before discussing the numerous DNN architectures widely used in geotechnical engineering, it is necessary to first discuss the computational tools and frameworks used to implement these algorithms. Regardless of the DL algorithm used, proper software platforms for designing, training and evaluating models are required. Significant progress has been achieved in the creation of such tools in recent years, with the introduction of a wide range of ML and DL frameworks that have been widely adopted by academics and developers across numerous scientific disciplines. However, certain frameworks have gained more traction within the civil engineering community, particularly in geotechnical engineering, due to their distinctive capabilities (Guan et al., 2023).
The nature of the problem, the preferred programming language, the volume and type of data and the overall study objectives all influence the tool selection. Table 2 summarizes the most important frameworks, including access links, primary and secondary programming languages, common application domains and related resources. Although these tools are relevant in a variety of industries, the domains shown in the table reflect the most common applications for the frameworks. Among them, TensorFlow, Keras and PyTorch have seen the highest adoption rates in geotechnical engineering research and will therefore be discussed in greater detail in the following sections.
Various frameworks for DL implementation
| No. | Framework names | Link to access | Primary programming language | Other programming language | Main application field* |
|---|---|---|---|---|---|
| 1 | TensorFlow | Link to tensorflowLink to the wbsite of tensorflow | Python | C++, Java, Go and Swift | CV, NLP, SR, RL |
| 2 | Keras | Link to kerasLink to the website of keras | Python | R | CV, NLP, RP |
| 3 | PyTorch | Link to pytorchLink to the website of pytorch | Python | C++, Java and Julia | CV, NLP, GM, RL |
| 4 | Caffe | Link to caffe.berkeleyvisionLink to the website of caffe.berkeleyvision | C++ | Python and MATLAB | CV, OD, S |
| 5 | MXNet | Link to mxnet.apacheLink to the website of mxnet.apache | – | Python, C++, R, Julia, Scala, Perl and MATLAB | CV, NLP, RS |
| 6 | Theano | Link to deeplearningLink to the website of deeplearning | Python | C and MATLAB | CV, NLP, SR |
| 7 | Torch | Link to torchLink to the website of torch | Lua | Python (PyTorch) and R (Torch7) | CV, NLP |
| 8 | Microsoft Cognitive Toolkit (CNTK) | Link to microsoftLink to the website of microsoft. | – | Python, C++ and C# | CV, SR, NLP |
| 9 | Chainer | Link to chainerLink to the website of chainer | Python | C++ | CV, NLP, RL |
| 10 | DeepLearning4j | Link to deeplearning4jLink to the website of deeplearning4j | Java and Scala | JVM languages like Kotlin and Clojure | FD, CBA |
| 11 | PaddlePaddle | Link to paddlepaddleLink to the website of paddlepaddle | Python and C++ | CV, NLP, RS |
| No. | Framework names | Link to access | Primary programming language | Other programming language | Main application field* |
|---|---|---|---|---|---|
| 1 | TensorFlow | Python | C++, Java, Go and Swift | CV, NLP, SR, | |
| 2 | Keras | Python | R | CV, NLP, | |
| 3 | PyTorch | Python | C++, Java and Julia | CV, NLP, GM, | |
| 4 | Caffe | C++ | Python and | CV, OD, S | |
| 5 | MXNet | – | Python, C++, R, Julia, Scala, Perl and | CV, NLP, | |
| 6 | Theano | Python | C and | CV, NLP, | |
| 7 | Torch | Lua | Python (PyTorch) and R (Torch7) | CV, | |
| 8 | Microsoft Cognitive Toolkit ( | – | Python, C++ and C# | CV, SR, | |
| 9 | Chainer | Python | C++ | CV, NLP, | |
| 10 | DeepLearning4j | Java and Scala | FD, | ||
| 11 | PaddlePaddle | Python and C++ | CV, NLP, |
* CV = Computer vision; NLP = natural language processing; SR = speech recognition; RL = reinforcement learning; RP = rapid prototyping; GM = generative models; OD = object detection; S = segmentation; RS = recommendation systems; FD = fraud detection; CBA = customer behavior analysis
3.1 TensorFlow
TensorFlow, a widely used open-source platform for numerical computation, is well-known for its contribution to ML endeavors, demonstrating adaptability across diverse processing devices such as CPUs, GPUs and TPUs. Its application is critical in geotechnical engineering, providing a flexible solution to a variety of problems (Abadi et al., 2016). In the realm of computer vision, TensorFlow makes it easier to recognize patterns and features in geotechnical imaging, such as soil sample cracks or changes in soil textures (Alzubaidi et al., 2021; De Araújo et al., 2024).
This capability is critical to improving our understanding and analysis of soil mechanics. Furthermore, TensorFlow’s generative modeling capabilities, notably with generative adversarial networks (GANs), allow for the creation of realistic three-dimensional models of geological formations (Goodfellow et al., 2014b; Song et al., 2021a; Song et al., 2021b; Song et al., 2020). These models are critical in applications such as reservoir modeling, where detailed representations of geological features are required.
Also, TensorFlow’s capabilities for reinforcement learning tasks create possibilities for optimizing the performance of geotechnical machinery such as tunnel boring machines (TBMs) through models that are trained to recognize optimal operating parameters (Jia et al., 2023; Liu and Yang, 2024; Liu et al., 2024c; Xuanyu et al., 2024). While Natural Language Processing (NLP) is not a common consideration in the realm of geotechnical engineering, the NLP functionalities supported by TensorFlow permit textual data analysis related to geotechnical projects in the form of project documentation and related textual information. This helps extract important details and better understand project data (Ding et al., 2022; Erfani and Cui, 2022; Locatelli et al., 2021; Corneli et al., 2023).
TensorFlow’s capability for time series analysis will be beneficial to forecasting geotechnical conditions with historical data. This function is especially useful for predicting variables, such as soil moisture content, assisting with better informed decisions. Furthermore, TensorFlow’s support of graph neural networks (GNNs) will also assist in working with complicated structures of data in geotechnical engineering. This aids in understanding the connections between geological properties to forecast how systems will behave using these under various conditions (Ahmed et al., 2021; Ferludin et al., 2022).
While TensorFlow offers several advantages, including a comprehensive ecosystem for ML development and integrated visualization tools such as TensorBoard (Abadi et al., 2016), it also presents certain limitations. GPU acceleration is primarily optimized for NVIDIA hardware through CUDA. Furthermore, TensorFlow’s extensive API and computational graph abstractions may result in a steeper learning curve compared with some competing frameworks. Previous studies have shown that training performance varies according to the neural network architecture and hardware configuration, with PyTorch outperforming TensorFlow in certain training scenarios (Dai et al., 2022).
3.2 Keras
Keras is an advanced neural network API that has been meticulously crafted in Python and supports multiple backends, including TensorFlow, JAX and PyTorch. This compatibility significantly streamlines the prototyping process, allowing for rapid and efficient development (Alom et al., 2019; Huang et al., 2015). Keras supports both convolutional and recurrent network architectures and ensures smooth computational transitions between CPU and GPU settings. Keras’ comprehensive versatility makes it an indispensable resource in the field of geotechnical engineering, where its wide range of functionalities can be applied to a variety of scenarios (Bergstra et al., 2010; Salvaris et al., 2018).
Specifically, Keras’ ability to implement CNNs plays a pivotal role in the analysis of imagery relevant to soil mechanics, allowing the identification of structural anomalies within soil samples or changes in their textural properties (Wang et al., 2022; Wu et al., 2021a; Zhang et al., 2021a; Yao et al., 2019). Keras also improves sequential data analysis by supporting RNNs, such as the specialized LSTM networks, which makes it easier to predict operational parameters for TBMs.
In addition to these specific applications, Keras holds considerable promise in soil type assessments and slope stability evaluation. The application of AI algorithms, facilitated through Keras, can be considered a new approach, affecting a considerable advancement in the classification of soil compared to older methods, which relied upon the interpretation of trained professionals and was prone to human error. Similarly, adopting an AI approach, which uses geological mapping, borehole logs and additional data through Keras to facilitate better slope stability assessments can improve landslide risk modeling and overall accuracy (Ronaldo, 2021).
The benefits of Keras include its intuitive, modular design, tailored for rapid experimentation or research around the use of deep chain neural networks. The user-friendly design and optimization for smaller data sets may appeal enough to attract developers, who like an easier, more fitting framework. However, Keras does not come without limitations, primarily an inability to adequately control complex manipulations of functions, as well as potential reported compatibility issues between Keras and TensorFlow code in its 2.0 version, which may also limit comprehensive modeling or detailed integration (Modi et al., 2024; Enawugaw and Yayeh, 2023; Arif et al., 2025).
3.3 PyTorch
PyTorch, an open-source ML library built on Torch, has gained popularity in a variety of fields, including NLP. It was primarily developed by Facebook’s AI Research lab and is a powerful tool for a wide range of tasks, including geotechnical engineering (Han et al., 2022; Valdes-Korovkin et al., 2024).
Geospatial DL is one of PyTorch’s most remarkable applications in geotechnical engineering. Practitioners can use the TorchGeo library, a PyTorch domain-specific library, to access data sets, samplers, transforms and pre-trained models for geospatial data (Stewart et al., 2022). These resources are extremely useful for tasks such as mapping land cover, monitoring deforestation and floods, tracking glacial flows, estimating hurricane intensity and detecting structures and roads. Furthermore, PyTorch’s capability in time series analysis provides predictive insights into future conditions using historical data, such as soil moisture levels, which is critical for making informed geotechnical engineering decisions (Zhang and Wang, 2024).
In addition, PyTorch Geometric expands PyTorch’s functionality to support DL on irregular input data such as graphs, point clouds and manifolds. This extension is especially useful for analyzing relationships between geological features and predicting the behavior of geotechnical systems based on component interactions (Cao et al., 2020). PyTorch also excels at image classification tasks, allowing for the classification of soil or rock types from images, as well as the implementation of generative models such as GANs for creating realistic 3D geological structure models, which are required in applications such as reservoir modeling (Zhu and Hu, 2024).
The benefits of using PyTorch in geotechnical engineering and beyond include its simplicity, ease of use, flexibility, efficient memory usage and the ability to generate dynamic computational graphs. Its data parallelism feature, which allows data to be distributed across multiple GPUs for processing, is highly efficient and popular among researchers (Akbar Firoozi and Firoozi, 2023; Pei and Qiu, 2024). However, as the newest addition to the DL framework landscape, PyTorch may face limitations in maturity and industrial adoption when compared to its predecessors, which were more widely used in academic circles (Nguyen et al., 2019).
4. Overview of common DL algorithms
Based on the bibliometric analysis in Section 2, some DL architectures clearly appear more often in geotechnical research than others. In particular, CNNs, LSTMs and GANs have proven to be especially effective and versatile for a variety of geotechnical problems. In Section 5, we take a closer look at these three methods, discussing their core principles, key advantages and examples of how they are applied in real geotechnical studies.
4.1 Convolutional neural network
CNNs are an important component of DL techniques, as they use the power of multiple layers to create efficient models (Krizhevsky et al., 2017). They are widely used in a variety of computer vision applications due to their efficacy. A CNN is typically made up of three primary layers: the convolutional layer, the pooling layer and the fully connected layer, each of which performs a distinct function. One of the primary advantages of CNNs is their ability to extract spatial features from data via their kernel (Cao et al., 2018; Jiang et al., 2023). For example, CNNs can recognize edges, color distributions and other spatial properties in images, making them extremely useful for image classification and other spatial data-related tasks.
When discussing CNNs, we frequently refer to two-dimensional CNNs, which are primarily used for image classification. However, one-dimensional and three-dimensional CNNs exist, each with their own set of real-world applications (Cao et al., 2018). The one-dimensional CNN (Conv1D) is used to process time-series data, such as that collected by a person’s accelerometer. This data, which includes acceleration values across three axes, can be analyzed to identify activities such as standing, walking and jumping. One-dimensional CNNs can also be used with audio and text data, which can be represented as time series (Muralidharan et al., 2021). The standard two-dimensional CNN (Conv2D), introduced in the LeNet-5 architecture, is primarily applied to image data. It is known as a two-dimensional CNN because the kernel moves across two dimensions of the data (Chen et al., 2021). Three-dimensional CNNs (Conv3D) are primarily used with 3D image data, such as Magnetic Resonance Imaging (MRI), which is widely used to examine the brain, spinal cord and internal organs. Conv3D can also be applied to video data, which is essentially a series of image frames with spatial features (Tiwari et al., 2023).
There are two main steps in training a CNN: feed-forward and back-propagation. During the feed-forward stage, the network receives the input image and generates an output based on the input data. The network output is then calculated. The network parameters are adjusted based on the calculated error rate, which is determined by comparing the network output with the correct answer using a loss function.
The next step is the back-propagation stage, which is based on the calculated error rate. At this point, the gradient of each parameter is calculated according to the chain rule, and all the parameters are adjusted based on their effect on the error in the network. With updated parameters, the next feed-forward phase can be run. This process is repeated several times, after which the network training is completed.
4.1.1 Core components of the CNN architecture
CNNs are made up of three types of layers (Liu and Feng, 2021): convolutional, pooling and fully connected (Figure 3):
The convolutional neural network architecture begins with stacked input images labelled Input. A highlighted region from the input connects by dotted lines to feature maps in Conv Layer 1. The feature maps then pass to Pooling Layer 1, followed by Conv Layer 2 and Pooling Layer 2. Dotted outlines show local regions transferred between successive feature maps. An ellipsis indicates additional intermediate processing before the network reaches a fully connected layer. The convolution and pooling stages are enclosed within a dashed region labelled Feature Extraction. The fully connected neural network is enclosed within a separate dashed region labelled Classification and contains several connected nodes arranged in layers, with ellipses indicating additional nodes. Four arrows lead from this network to output boxes labelled Class 1, Class 2, Class 3 and Class 4.Architecture of a convolutional neural network (CNN)
The convolutional neural network architecture begins with stacked input images labelled Input. A highlighted region from the input connects by dotted lines to feature maps in Conv Layer 1. The feature maps then pass to Pooling Layer 1, followed by Conv Layer 2 and Pooling Layer 2. Dotted outlines show local regions transferred between successive feature maps. An ellipsis indicates additional intermediate processing before the network reaches a fully connected layer. The convolution and pooling stages are enclosed within a dashed region labelled Feature Extraction. The fully connected neural network is enclosed within a separate dashed region labelled Classification and contains several connected nodes arranged in layers, with ellipses indicating additional nodes. Four arrows lead from this network to output boxes labelled Class 1, Class 2, Class 3 and Class 4.Architecture of a convolutional neural network (CNN)
Convolutional layer: This is the fundamental layer of a CNN. The layer’s parameters are made up of a series of learnable filters or kernels that have a small receptive field but extend the entire depth of the input volume. During the forward pass, each filter is convolved across the width and height of the input volume, computing the dot product of the filter’s entries and the input, resulting in a two-dimensional activation map. The network learns filters that activate when it detects a visual feature such as an edge of some orientation or a blotch of some color on the first layer, and eventually entire honeycomb or wheel-like patterns on higher layers of the network.
Pooling layer: The pooling layer progressively reduces the spatial size of the representation to reduce the amount of parameters and computation in the network. This layer operates independently on every depth slice of the input and resizes it spatially. The most common approach used in pooling is max pooling, where the maximum value inside the window becomes the new value for the pooled output.
Fully connected layer: Fully connected layers connect every neuron in one layer to every neuron in another layer. It is the same as a traditional multilayer perceptron neural network (MLP). The flattened matrix goes through a fully connected layer to classify the images.
In addition to these layers, other important components include the dropout layer and the activation function. The dropout layer is a technique for preventing a model from overfitting. During training, it randomly drops (or temporarily removes) a number of the layer’s output features. The “dropout rate” is the percentage of features that are zeroed out, and it is typically set between 0.2 and 0.5. The activation function is used to add nonlinearity to the output of a neuron. This enables the model to learn more complex functions. Common activation functions include ReLU, tanh and the sigmoid function (Park and Kwak, 2017).
4.1.2 Advanced architectures built on CNNs
Several CNN architectures have been developed over the years, each with its unique configuration and contributions to the field of DL. Here are some of the most notable ones:
LeNet-5 (1998): Developed by Yann Lecun (Lecun et al., 1998), LeNet-5 is a pioneering seven-layer CNN that was primarily used for handwriting and character recognition. It was one of the first successful applications of CNNs. The architecture includes two sets of convolutional and average pooling layers, followed by a fully connected layer, and finally, a softmax classifier. It achieved 99.2% accuracy on isolated character recognition.
AlexNet (2012): AlexNet (Krizhevsky et al., 2017) is a deeper and broader version of LeNet that won the ImageNet Large-Scale Visual Recognition Challenge (ILSVRC) by a widelarge margin in 2012. It was developed by Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton. The network has 60 million parameters and 650,000 neurons. It is made up of five convolutional layers, some of which are followed by max-pooling layers, and three fully connected layers, culminating in a 1000-way softmax. To reduce overfitting, the authors used two regularization techniques: data augmentation and dropout.
VGGNet (2014): The VGGNet (Simonyan and Zisserman, 2015) was introduced by the Visual Geometry Group (VGG) from the University of Oxford, and it was a breakthrough in terms of architecture design. The most unique thing about VGGNet is that instead of having a large number of hyper-parameters, it focuses on having convolutional layers of 3 x 3 filter with a stride 1 and always uses SAME padding and maxpool layer of 2 x 2 filter of stride 2. It follows this arrangement of convolution and max pool layers consistently throughout the whole architecture. In total, the VGGNet consists of 16 convolutional layers and is very appealing because of its very uniform architecture.
GoogLeNet (2014): The ILSVRC 2014 winner was a convolutional network from (Szegedy et al., 2015). Its main contribution was the development of an inception module that dramatically reduced the number of parameters in the network (4M, compared to AlexNet with 60M). Additionally, this network used average pooling instead of fully connected layers at the top of the ConvNet, eliminating a large amount of parameters that do not seem to matter much. There are 22 layers in GoogLeNet.
ResNet (2015): ResNet (He et al., 2016), short for residual networks, is a classic neural network that serves as the foundation for many computer vision tasks. The main innovation of ResNet is the introduction of “skip connections,” which allow the gradient to be directly backpropagated to earlier layers. The central idea of ResNet is to introduce a so-called “identity shortcut connection” that skips one or more layers. The authors argue that stacking layers should not degrade network performance because identity mappings (layers that do nothing) could be stacked on top of the current network, and the resulting architecture would perform similarly. This indicates that the deeper model should not produce a training error higher than its shallower counterparts. They proposed a residual learning framework to ease the training of networks that are substantially deeper than those used previously, and provided comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth.
A summary of these architectures is provided in Table 3. Each of these architectures has significantly contributed to the field of DL, enabling the development of more complex and accurate models. They have been used in a variety of applications, such as image and video recognition, recommendation systems, image generation and many more.
Evolution of key CNN architectures and their features
| Year | Architecture | Developers | Key features |
|---|---|---|---|
| 1998 | LeNet-5 | Yann Lecun | 7-layer CNN for handwriting and character recognition, includes convolutional, average pooling, fully connected layers and a softmax classifier |
| 2012 | AlexNet | Alex Krizhevsky; Ilya Sutskever; Geoffrey Hinton. Research | Deeper and wider than LeNet, uses data augmentation and dropout for overfitting reduction, consists of convolutional layers, max-pooling, fully connected layers and softmax |
| 2014 | VGGNet | Visual Geometry Group; University of Oxford | Uniform architecture with 3 x 3 convolutional layers, stride 1, SAME padding and 2 x 2 max-pooling. Contains 16 convolutional layers |
| 2014 | GoogLeNet | Szegedy et al., 2015 | Introduced Inception Module, reduced parameters, replaced fully connected layers with average pooling, total 22 layers |
| 2015 | ResNet | He et al. | Introduced “skip connections” and “identity shortcut Connection,” facilitating direct gradient backpropagation and Supporting increased depth without performance degradation |
| Year | Architecture | Developers | Key features |
|---|---|---|---|
| 1998 | LeNet-5 | Yann Lecun | 7-layer |
| 2012 | AlexNet | Alex Krizhevsky; Ilya Sutskever; Geoffrey Hinton. Research | Deeper and wider than LeNet, uses data augmentation and dropout for overfitting reduction, consists of convolutional layers, max-pooling, fully connected layers and softmax |
| 2014 | VGGNet | Visual Geometry Group; University of Oxford | Uniform architecture with 3 x 3 convolutional layers, stride 1, |
| 2014 | GoogLeNet | Introduced Inception Module, reduced parameters, replaced fully connected layers with average pooling, total 22 layers | |
| 2015 | ResNet | He et al. | Introduced “skip connections” and “identity shortcut Connection,” facilitating direct gradient backpropagation and Supporting increased depth without performance degradation |
4.2 LSTM-based methods
Long short-term memory (LSTM) networks are a subset of recurrent neural networks (RNNs) that are specifically designed to model temporal sequences and capture long-range dependencies in time series or sequential data. LSTMs, initially proposed by Hochreiter and Schmidhuber (Hochreiter and Schmidhuber, 1997), address a significant limitation of conventional RNNs, namely the vanishing and exploding gradient problems, that impede effective learning of long-term patterns across lengthy sequences. To address this, LSTMs use a sophisticated architecture that includes memory cells and three types of gating units: input, forget and output gates (Al-Selwi et al., 2024).
These components work together to regulate the flow of information, allowing the network to maintain, update and discard data in a controlled manner over time. This dynamic gating mechanism enables LSTMs to maintain context over long time steps, making them particularly useful in a variety of applications including NLP, speech recognition, time-series forecasting and sequential decision-making tasks. Because of their ability to handle temporal dependencies, LSTM networks have become an essential tool in the development of DL models for sequential data analysis.
4.2.1 Core components of the LSTM architecture
A typical LSTM unit is composed of four fundamental components, each playing a crucial role in managing the flow of information through the network (Sherstinsky, 2020). These components include:
Input gate: Controls how much of the new input to incorporate into the memory cell.
Forget gate: Decides what information to discard from the cell state.
Cell state: Acts as a memory pathway, carrying useful information across time steps.
Output gate: Regulates what part of the memory cell should be output at each step.
These components work together to preserve and propagate relevant information while filtering out noise, making LSTMs highly suitable for time-dependent tasks (Figure 4).
The schematic presents an L S T M memory cell enclosed by a dashed boundary. The previous cell state, C subscript t minus 1, enters from the left along the upper horizontal path. At the forget gate, a sigma function receives the previous hidden state, h subscript t minus 1, and the current input, X subscript t. Its output passes to a multiplication node on the cell state path. The input gate contains a sigma function and a hyperbolic tangent function supplied by h subscript t minus 1 and X subscript t. Their outputs meet at a multiplication node, then pass upward to an addition node, where they combine with the retained previous cell state. The updated cell state continues right as C subscript t. At the output gate, the updated cell state passes through a hyperbolic tangent function. A sigma function receives h subscript t minus 1 and X subscript t, and its output multiplies the transformed cell state. The result exits to the right as h subscript t and also upwards as h subscript t. Arrows indicate the direction of information flow through all gates and operations.Architecture of a LSTM
The schematic presents an L S T M memory cell enclosed by a dashed boundary. The previous cell state, C subscript t minus 1, enters from the left along the upper horizontal path. At the forget gate, a sigma function receives the previous hidden state, h subscript t minus 1, and the current input, X subscript t. Its output passes to a multiplication node on the cell state path. The input gate contains a sigma function and a hyperbolic tangent function supplied by h subscript t minus 1 and X subscript t. Their outputs meet at a multiplication node, then pass upward to an addition node, where they combine with the retained previous cell state. The updated cell state continues right as C subscript t. At the output gate, the updated cell state passes through a hyperbolic tangent function. A sigma function receives h subscript t minus 1 and X subscript t, and its output multiplies the transformed cell state. The result exits to the right as h subscript t and also upwards as h subscript t. Arrows indicate the direction of information flow through all gates and operations.Architecture of a LSTM
Several architectural variations of the standard LSTM network have been developed over time to enhance its performance and adapt it to specific types of data and tasks. The basic LSTM layer processes sequences in a unidirectional manner, where each output is influenced by prior computations in the sequence. While this structure is effective for many temporal modeling tasks, it can be limited when future context is also informative (Liu et al., 2024a; Wang et al., 2023a). To address this, the Bidirectional LSTM (BiLSTM) architecture was introduced, which consists of two separate LSTM layers processing the input sequence in forward and backward directions. This configuration enables the network to access both past and future contextual information simultaneously, significantly improving performance in tasks such as speech recognition and text analysis (Wu et al., 2024).
For tasks requiring more complex pattern extraction, the Stacked or Deep LSTM architecture is used. In this design, multiple LSTM layers are stacked vertically, with the output of one layer serving as the input to the next. This hierarchical structure allows the network to capture higher-level temporal abstractions and hierarchical dependencies in the data. In domains where both spatial and temporal correlations are crucial, such as precipitation forecasting or video frame prediction, the convolutional LSTM (ConvLSTM) becomes particularly valuable. ConvLSTM replaces standard matrix multiplications within the LSTM cell with convolutional operations, allowing it to preserve spatial relationships while modeling temporal dynamics (Shen et al., 2021; Mojtahedi et al., 2025).
Furthermore, LSTM performance can be significantly enhanced through the integration of attention mechanisms, leading to the development of attention-based LSTMs. These models enable the network to dynamically focus on the most relevant segments of the input sequence at each time step, improving interpretability and performance in tasks involving long sequences or complex dependencies, such as machine translation and document summarization. Finally, to enable the training of deeper LSTM networks without encountering issues such as vanishing gradients or performance degradation, the residual LSTM architecture incorporates skip connections between layers, inspired by the success of residual learning in CNNs (e.g. ResNet). These residual connections enable the model to retain information across layers more effectively, resulting in improved convergence and overall accuracy in deep sequence models (Liu et al., 2021; Liu and Feng, 2021).
4.2.2 Advanced architectures built on LSTM
Several advanced neural network architectures have been proposed and refined in recent years to address traditional LSTM limitations such as modeling long-term dependencies, limited parallelization and capturing complex spatial-temporal patterns. These advanced designs aim to improve the performance, representational depth and generalization capabilities of LSTM-based models for a variety of sequence modeling tasks:
Seq2Seq (Sequence-to-sequence) models: These models use an encoder-decoder architecture built on LSTMs to transform one sequence into another, such as converting a sentence in one language to another or summarizing long text into shorter forms. The encoder processes the entire input sequence into a fixed-length context vector, and the decoder uses this vector to generate the target sequence, one step at a time. This framework allows flexible input-output lengths and forms the foundation for many NLP tasks.
Hybrid CNN-LSTM networks: These architectures combine convolutional neural networks (CNNs) and LSTM layers to exploit both spatial and temporal data representations. CNNs capture local spatial patterns (e.g. in images, frames or sensor arrays), whereas LSTMs model sequential relationships over time. This hybrid design is especially effective in domains where data exhibits both spatial and temporal dynamics, such as video classification, motion analysis and multivariate sensor data modeling.
Transformer-LSTM hybrids: This approach combines the strengths of Transformer models − particularly their self-attention mechanism that captures long-range dependencies efficiently − with the sequential processing capabilities of LSTMs. In such hybrids, Transformers often provide a global view of the input sequence, while LSTMs refine this information with their ability to model fine-grained temporal dependencies. These architectures are well-suited for tasks like sequence forecasting, time-series modeling and language understanding, where both global context and temporal structure are crucial.
4.3 GAN-based methods
GANs, first proposed by Ian Goodfellow et al. (Goodfellow et al., 2014a), are a pioneering class of DL models for generative tasks. GANs are based on a game-theoretic framework that involves training two competing neural networks: a generator (G) and a discriminator (D). The generator is responsible for creating synthetic data samples that closely resemble the real data distribution, while the discriminator determines whether each sample is real or generated. During training, the two networks compete in a minimax game, with the generator attempting to fool the discriminator and the discriminator attempting to correctly classify inputs.
Through this adversarial process, both models iteratively improve, eventually allowing the generator to produce data that is increasingly similar to real-world examples. GANs have made significant advances in the field of generative modeling, with far-reaching implications for computer vision, image synthesis, data augmentation and, more recently, scientific simulations. Their ability to learn complex data distributions without explicit modeling has created new opportunities for innovative applications and realistic data generation in academic and industrial settings (Wang et al., 2023b).
4.3.1 Core components of the GAN architecture
The GAN architecture (Zhang et al., 2021b; Zhang et al., 2024) consists of three main components that work together to create realistic data samples (Figure 5):
The generative adversarial network framework contains a Generator A N N and a Discriminator A N N. Random Noise enters the Generator A N N from the left. The generator produces a stack labelled Generated Fake Samples, which is passed to the discriminator. Training Data enters a stack labelled Real Samples, which is also passed to the discriminator. The discriminator receives both real and generated samples and evaluates them. Two output paths from the discriminator lead to Discriminator Loss and Generator Loss on the right. Feedback arrows from the loss paths return towards the discriminator, while a lower feedback path from the discriminator returns to the generator. The labels Real Samples and Generated Fake Samples appear above and below their respective dashed sample groups.Architecture of a GAN
The generative adversarial network framework contains a Generator A N N and a Discriminator A N N. Random Noise enters the Generator A N N from the left. The generator produces a stack labelled Generated Fake Samples, which is passed to the discriminator. Training Data enters a stack labelled Real Samples, which is also passed to the discriminator. The discriminator receives both real and generated samples and evaluates them. Two output paths from the discriminator lead to Discriminator Loss and Generator Loss on the right. Feedback arrows from the loss paths return towards the discriminator, while a lower feedback path from the discriminator returns to the generator. The labels Real Samples and Generated Fake Samples appear above and below their respective dashed sample groups.Architecture of a GAN
Generator: Typically a neural network (often fully connected or convolutional) that takes random noise (latent vector) and produces synthetic samples.
Discriminator: A classifier that evaluates whether a given sample is real (from the data set) or fake (from the generator).
Adversarial loss: A minimax objective function used to optimize the generator and discriminator in a zero-sum game.
GANs have evolved with the incorporation of various specialized layers to enhance their performance across different tasks. Each type of layer serves a specific purpose in improving the network’s ability to generate high-quality data.
Initially, fully connected layers were used in early GAN models, such as vanilla GANs, in which each neuron in one layer is linked to every neuron in the next. However, fully connected layers are ineffective for high-dimensional data such as images because they do not capture spatial relationships within the data. As GANs evolved, convolutional layers became the foundation of modern architectures, particularly deep convolutional GANs (DCGANs) (Campos Montero et al., 2025). These layers are especially useful for image generation because they are designed to capture local patterns and spatial hierarchies, making them essential for processing visual data.
The generator network incorporates transposed convolutional layers (or deconvolutional layers) to improve the generation of high-resolution outputs. These layers convert the low-dimensional noise input into detailed, high-resolution data. This process enables the generator to produce images with finer details, which improves the quality of the generated samples (Baghbani et al., 2022). Training stability and efficient convergence are common challenges in GANs, and normalization layers such as batch normalization, layer normalization and spectral normalization are frequently used to address these issues. These layers normalize the activations in the network, reducing issues like vanishing or exploding gradients and allowing the network to train more efficiently (Jiang et al., 2023). Deeper GAN architectures incorporate residual blocks to improve gradient flow and overall performance. These blocks allow the network to skip certain layers and add their outputs to subsequent layers, reducing performance degradation as the network grows deeper and more complex.
Some advanced GAN models, such as attention GANs or self-attention GANs (SAGANs), include self-attention layers. These layers enable the model to focus on specific, important regions of the input data, capturing long-term dependencies that are essential for producing high-quality data. This mechanism is especially useful for producing more realistic and consistent results (Wu et al., 2023).
4.3.2 Advanced architectures built on GAN
Over the years, numerous variants and extensions of the original GAN model have been proposed to address challenges like mode collapse, training instability and limited diversity. Some of the most important GAN variants include:
DCGAN (deep convolutional GAN):DCGAN (Patil, 2021; Liu et al., 2023a) added convolutional layers to GAN architectures, making it one of the most significant advances in GAN research. DCGANs are more effective for image processing because they use convolutional layers instead of fully connected ones. The convolutional structures enable the network to detect spatial hierarchies and local patterns in images, significantly improving the quality and realism of the generated images. DCGAN became the standard model for image generation tasks, demonstrating GANs’ power in visual data synthesis.
Conditional GAN (cGAN): Conditional GANs (Chrysos et al., 2020) build on the original GAN framework, incorporating additional information or conditions into the generation process. This auxiliary data, such as class labels or input images, serves as an additional input to both the generator and the discriminator. cGANs are especially useful for tasks requiring precise generation, such as image-to-image translation or label-conditioned synthesis. For example, a cGAN can generate an image of a specific object or scene based on a class label, making it an effective tool for image editing, colorization and super-resolution.
CycleGAN: CycleGAN (Zhu et al., 2017) addresses the challenge of image-to-image translation when paired training data is unavailable. Traditional GANs require paired data sets (e.g. a set of images of horses and a corresponding set of images of zebras), but CycleGAN can perform unpaired image translation. This makes it particularly useful for tasks like converting aerial images to maps, or transforming sketches into photographs. CycleGAN works by learning a mapping between two domains (e.g. photos and sketches) while preserving important features of the original images through the use of cycle consistency, which ensures that a translated image can be reverted back to its original form.
Pix2Pix: Pix2Pix (Isola et al., 2017) is a supervised image-to-image translation model that trains a GAN on paired data sets to perform tasks such as converting satellite images to maps and edge images to photographs. The model is trained on a set of paired examples (for example, a photo of a landscape and its corresponding edge map) and learns to produce the desired output. Unlike CycleGAN, Pix2Pix requires paired training data and is especially useful for tasks requiring precise mapping between input and output images. It uses a conditional GAN framework in which the generator learns to map input images to output images, and the discriminator assesses the realism of the generated results.
Wasserstein GAN (WGAN) and WGAN-GP: Wasserstein GANs (Arjovsky et al., 2017; Gulrajani et al., 2017; Shahrestani, 2024) modify the loss function to address some of the challenges associated with GAN training, such as instability and mode collapse. Instead of the traditional binary cross-entropy loss, WGAN uses the Earth-Mover (Wasserstein) distance, which quantifies the disparity between the distributions of real and generated data. This new loss function improves training stability and enables the model to detect when the generator fails to capture the actual data distribution. The WGAN-GP variant improves training by including a gradient penalty term, which prevents the discriminator from overfitting and contributes to better convergence.
StyleGAN & StyleGAN2: StyleGAN (Karras et al., 2019) and its successor StyleGAN2 pioneered a novel method for controlling the style and features of generated images. These models enable style-based control at multiple levels, ranging from coarse features (such as the overall shape of a face) to fine details (such as skin texture). This makes StyleGANs especially useful for applications like highly realistic human face synthesis, where specific attributes (e.g. age, hairstyle or expression) can be changed while image fidelity remains high. StyleGAN and StyleGAN2 are widely used in the creative industries and research to generate realistic, editable images.
3D-GAN & VoxelGAN: 3D-GANs and VoxelGANs (Khan et al., 2023) extend the original GAN model to handle volumetric and 3D data. Unlike traditional 2D image generation, these models are capable of generating 3D objects or structures, such as models used in scientific simulations or computer-aided design (CAD). They work by generating 3D voxel grids, which are 3D representations of an object in the form of a grid of cubes (voxels). These models have been used in applications such as 3D medical imaging, object recognition and the design of virtual environments. By incorporating 3D data generation into the GAN framework, these models enable a new dimension of generative modeling in fields like virtual reality and scientific modeling.
4.4 Emerging DL architectures
In recent years, several advanced ML and DL architectures have emerged beyond conventional CNNs, LSTMs and GANs. These methods aim to address limitations related to long-range dependency modeling, irregular data structures, incorporation of physical knowledge, computational efficiency and model interpretability. Notable developments include Transformer architectures based on self-attention mechanisms (Vaswani et al., 2017), graph neural networks (GNNs) for graph-structured data (Wu et al., 2021b), physics-informed neural networks (PINNs) that integrate governing physical laws into the learning process (Raissi et al., 2019), neural operator learning frameworks such as DeepONets (Lu et al., 2021) and explainable artificial intelligence (XAI) techniques that improve model transparency and interpretability (Lundberg and Lee, 2017). A brief overview of these developments is provided in Table 4.
Summary of emerging machine learning and deep learning architectures, including their core concepts and key advantages
| Architecture | Core concept | Key advantages |
|---|---|---|
| Transformers / attention | Uses self-attention mechanisms to capture long-range dependencies and global context within data | Effective for modeling complex sequential and spatial relationships; scalable and flexible across data modalities |
| Graph neural networks (GNNs) | Processes data represented as graphs, where nodes and edges encode relationships among entities | Naturally handles irregular and non-Euclidean data while capturing complex relational structures |
| Physics-Informed neural networks (PINNs) | Incorporates governing physical laws and constraints into the learning process through the loss function | Improves physical consistency, interpretability and generalization, particularly when data are limited |
| Operator learning methods (DeepONet, FNO) | Learns mappings between function spaces rather than discrete input-output pairs | Enables rapid approximation of complex physical systems and efficient surrogate modeling |
| Explainable AI (XAI) | Provides methods to interpret and explain model predictions and decision-making processes | Improves transparency, trustworthiness and validation of machine learning models |
| Architecture | Core concept | Key advantages |
|---|---|---|
| Transformers / attention | Uses self-attention mechanisms to capture long-range dependencies and global context within data | Effective for modeling complex sequential and spatial relationships; scalable and flexible across data modalities |
| Graph neural networks (GNNs) | Processes data represented as graphs, where nodes and edges encode relationships among entities | Naturally handles irregular and non-Euclidean data while capturing complex relational structures |
| Physics-Informed neural networks (PINNs) | Incorporates governing physical laws and constraints into the learning process through the loss function | Improves physical consistency, interpretability and generalization, particularly when data are limited |
| Operator learning methods (DeepONet, | Learns mappings between function spaces rather than discrete input-output pairs | Enables rapid approximation of complex physical systems and efficient surrogate modeling |
| Explainable | Provides methods to interpret and explain model predictions and decision-making processes | Improves transparency, trustworthiness and validation of machine learning models |
5. Transfer learning
Transfer learning, while not a distinct neural network architecture, has become a popular DL strategy, thanks to its ability to improve performance when domain-specific data is limited. This method uses a pre-trained DL model that has already been trained on large-scale data sets, typically from other domains such as computer vision or NLP. The model is then fine-tuned for a particular task in a new, usually smaller, data set or domain, such as geotechnical engineering (Liu et al., 2024b; Gao, 2024).
The core concept of transfer learning is based on the assumption that knowledge learned in one task or domain can be transferred and reused in another, as long as the tasks have similar underlying patterns. Transfer learning typically involves several steps:
Choosing a pre-trained model: The first step is selecting a model that has already been trained on large benchmark data sets, such as ImageNet for images or COCO for object detection. These models capture a wide range of general features that can be useful across different tasks.
Freezing the lower layers: In the next step, the lower layers of the pre-trained model, which capture general features like edges and textures, are frozen. These layers are already trained to detect fundamental patterns that are likely to be useful for a wide range of tasks, so there is no need to retrain them.
Fine-tuning the upper layers: The pre-trained model’s upper layers, which are task-specific, are then fine-tuned. Alternatively, custom layers can be added to the model to better tailor it to the intended task. For example, in geotechnical engineering, this could include tasks such as soil classification or slope stability prediction, for which the model would be trained on a smaller, domain-specific data set.
Training the modified model: Finally, the modified model is trained on a smaller geotechnical data set to optimize it for the task at hand. This fine-tuning process enables the model to learn the nuances of the new domain while retaining the general features it has learned from the large-scale data set.
The workflow of this process is schematically shown in Figure 6, illustrating how a pre-trained model can be adapted to a specific task in geotechnical engineering through transfer learning. By reusing and refining the knowledge learned from large-scale data sets, transfer learning enables DL models to perform effectively even in domains with limited data, such as geotechnical engineering.
The diagram compares a Pretrained Model in the upper row with a Custom Model in the lower row. In both rows, an Input passes through four sequential feature-processing layers connected by right-pointing arrows. A horizontal bracket labelled Common Inner Layer spans these shared layers, indicating that they are reused in both models. In the pretrained model, the final feature layer connects to a fully connected network containing several node layers and ellipses for additional nodes, ending in an Output with multiple output nodes. In the custom model, the same common inner layers feed a newly configured fully connected network. A bracket labelled Custom Final Layers spans the custom network and its output section. The custom model ends with its own Output containing multiple output nodes.Architecture of a transfer learning
The diagram compares a Pretrained Model in the upper row with a Custom Model in the lower row. In both rows, an Input passes through four sequential feature-processing layers connected by right-pointing arrows. A horizontal bracket labelled Common Inner Layer spans these shared layers, indicating that they are reused in both models. In the pretrained model, the final feature layer connects to a fully connected network containing several node layers and ellipses for additional nodes, ending in an Output with multiple output nodes. In the custom model, the same common inner layers feed a newly configured fully connected network. A bracket labelled Custom Final Layers spans the custom network and its output section. The custom model ends with its own Output containing multiple output nodes.Architecture of a transfer learning
Transfer learning offers several significant benefits that make it an attractive approach, especially when working with limited domain-specific data. One of the primary advantages is the reduced data requirements, as transfer learning enables the training of accurate models even with relatively small data sets. By leveraging pre-trained models, training times are also faster, as the pre-trained weights significantly accelerate the convergence process, reducing the time needed for training from scratch (Thiruchittampalam et al., 2024). Furthermore, models that have been pre-trained on diverse data sets exhibit better generalization, making them more robust and capable of performing well across a wide range of tasks and domains. Furthermore, transfer learning gives researchers access to cutting-edge architectures, allowing them to use well-tested models such as ResNet, VGG, Inception, BERT and even pre-trained LSTM and Transformer models. This access to advanced, pre-existing models enables the use of cutting-edge techniques in new applications, improving the overall performance of the trained models (Fang et al., 2023).
Transfer learning has shown great promise in several domains of geotechnical engineering, providing innovative solutions to a variety of challenges. One notable application is soil and rock property prediction, where transfer learning can be used to fine-tune vision-based models using borehole image logs or laboratory photos. This method enables practitioners to estimate critical parameters such as soil classification, grain size distribution and rock type even with limited labeled data. Pre-trained models in slope stability analysis can be modified to identify stable versus unstable slopes using image-based data, such as drone or satellite images, or numerical simulation outputs. Transfer learning also plays a key role in remote sensing and subsurface mapping, where models trained on global satellite data sets can be fine-tuned for high-resolution, local geotechnical mapping tasks like landslide detection or ground deformation monitoring. Furthermore, synthetic data and simulation integration benefits from transfer learning by enhancing model generalization when combining real-world measurements with simulated data, such as those obtained from finite element analysis or discrete element modeling. This hybrid approach allows for the creation of more robust physics-data models. Additionally, cross-domain knowledge transfer is increasingly valuable, as knowledge gained from structural or transportation engineering data sets can be applied to geotechnical tasks involving structural interactions, such as predicting the performance of retaining walls. These applications highlight how transfer learning can bridge the gap between different data sets and domains, improving accuracy and efficiency in geotechnical engineering tasks (Zhang et al., 2021b; Tao et al., 2025).
6. Data sources
Data is the primary input for deep neural networks. Regardless of how much effort is put into optimizing network structure and architecture, a DL model’s performance will suffer significantly if the training data is insufficient or inaccurate. As a result, identifying reliable data sources and using appropriate preprocessing techniques are critical components of establishing DNNs effectively. This section discusses the most common data sources used in geotechnical engineering (Phoon, 2020).
Geotechnical engineers collect data from a variety of sources, the most important of which are laboratory testing, field measurements, remote sensing and Geographic Information Systems (GIS)-based data sets. Historically, geotechnical engineers relied heavily on laboratory tests to evaluate soil and rock properties. These tests produce structured data, including grain size distribution, density, shear strength parameters and hydraulic properties. To ensure consistency and reliability for DL model training, laboratory data must be quality controlled, cleaned and, in some cases, normalized (Onyelowe et al., 2023; Parsa-Pajouh, 2025).
Due to several limitations associated with laboratory testing, such as sampling difficulties, sample disturbance and transportation issues, field testing is also commonly used in geotechnical engineering. Geotechnical instruments like piezometers, inclinometers and settlement plates gather information from field sites and subsurface layers. This real-time monitoring provides valuable information about how soil and rock behave under different conditions. Filtering, noise reduction and synchronization are common techniques used to preprocess field data before it is used in DL applications (Yousefpour et al., 2024).
Given the vast scale of geotechnical sites, physical investigation alone cannot provide all of the necessary information. As a result, geotechnical studies now rely heavily on remote sensing technologies such as satellite imagery, LiDAR and UAVs. These data sources provide a more comprehensive understanding of geological and environmental characteristics. To prepare such data for neural network training, image processing methods such as georeferencing, resizing, color-space transformation and feature extraction are required (Mohan et al., 2021).
With the advancement of GIS, GIS-based data has become widely used in geotechnical engineering. These data sets may contain a variety of layers such as land use, topography, geological maps and hydrological data, which can be combined to create comprehensive geotechnical data sets. Data fusion, interpretation and spatial analysis are examples of preprocessing techniques used to prepare GIS data for DL applications. The combination of GIS and DL improves the accuracy of geotechnical predictions and output visualization (Raja et al., 2023; Hassan et al., 2024; Kim and Ji, 2022).
Publicly available data sets are another important data source that is used in many disciplines, including geotechnical engineering. These include data published by government agencies, research institutions and academic journals. Due to a lack of awareness about available geotechnical data sets, researchers frequently incur unnecessary financial and time costs in data collection. Thus, familiarity with the major data repositories is essential. Table 5 shows a selection of key geotechnical data sources and their characteristics.
Geotechnical data repositories and access information
| No. | Name of source | Description | Access link |
|---|---|---|---|
| 1 | Kaggle data sets | General data repository; includes geotechnical-related data sets (soil, slope, foundation data) by search | Link to kaggleLink to the website of kaggle |
| 2 | Data in brief (Elsevier journal) | A peer-reviewed journal publishing standalone data sets across disciplines, including geotechnical data sets. Authors sometimes upload complete lab/field data | Link to sciencedirectLink to the website of sciencedirect |
| 3 | Harvard Dataverse | Open repository for academic data sets. Geotechnical and civil engineering data sets can be found by searching relevant keywords (e.g. “soil,” “foundation,” “CPT”) | Link to dataverse.harvardLink to the website of dataverse.harvard |
| 4 | National Geotechnical Experimentation Sites (NGES) database | Contains a wealth of information on soil properties, rock mechanics and other geotechnical data gathered from field sites all over the world. Researchers and engineers can use this information to better understand the behavior of soils and rocks under various conditions | Link to geotechdataLink to the website of geotechdata |
| 5 | USGS Earthquake Hazards Program | Provides data on seismic activity, ground shaking and soil liquefaction potential, which are important considerations in geotechnical engineering for earthquake-prone regions. Engineers can access this data to assess seismic hazards and design resilient structures | Link to earthquakeLink to the website of earthquake |
| 6 | Geotechnical Engineering Database (GEDB) | A repository of geotechnical data, case histories and research papers curated by the International Society for Soil Mechanics and Geotechnical Engineering (ISSMGE). Engineers can access this data for reference and research purposes | Link to geotechdataLink to the website of geotechdata |
| 7 | British Geological Survey (BGS) | Provides geotechnical data, maps and reports on the geology and ground conditions in the UK. Engineers can access this data for site investigations and geological hazard assessments | Link to bgsLink to the website of bgs |
| 8 | GEER Data Portal | Post-event geotechnical reconnaissance data (earthquake, landslide, flood, etc.) | Link to datacenterhubLink to the website of datacenterhub |
| 9 | DesignSafe-CI | Repository of published geotechnical and structural engineering data (NSF-funded projects) | Link to designsafe-ciLink to the website of designsafe-ci |
| No. | Name of source | Description | Access link |
|---|---|---|---|
| 1 | Kaggle data sets | General data repository; includes geotechnical-related data sets (soil, slope, foundation data) by search | |
| 2 | Data in brief (Elsevier journal) | A peer-reviewed journal publishing standalone data sets across disciplines, including geotechnical data sets. Authors sometimes upload complete lab/field data | |
| 3 | Harvard Dataverse | Open repository for academic data sets. Geotechnical and civil engineering data sets can be found by searching relevant keywords (e.g. “soil,” “foundation,” “CPT”) | |
| 4 | National Geotechnical Experimentation Sites ( | Contains a wealth of information on soil properties, rock mechanics and other geotechnical data gathered from field sites all over the world. Researchers and engineers can use this information to better understand the behavior of soils and rocks under various conditions | |
| 5 | Provides data on seismic activity, ground shaking and soil liquefaction potential, which are important considerations in geotechnical engineering for earthquake-prone regions. Engineers can access this data to assess seismic hazards and design resilient structures | ||
| 6 | Geotechnical Engineering Database ( | A repository of geotechnical data, case histories and research papers curated by the International Society for Soil Mechanics and Geotechnical Engineering ( | |
| 7 | British Geological Survey ( | Provides geotechnical data, maps and reports on the geology and ground conditions in the | |
| 8 | Post-event geotechnical reconnaissance data (earthquake, landslide, flood, etc.) | ||
| 9 | DesignSafe-CI | Repository of published geotechnical and structural engineering data (NSF-funded projects) |
7. Conclusion
This paper provides a comprehensive review of DL methodologies in geotechnical engineering, with the primary goal of cataloging, categorizing and explaining the most commonly used algorithms and computational frameworks. Recognizing the growing importance of DL in addressing complex geotechnical challenges, this study focuses on the fundamental tools and techniques required for researchers and practitioners entering this interdisciplinary field.
The principal contributions of this review are as follows:
An overview of common computational frameworks (e.g. TensorFlow, Keras and PyTorch), including their roles and typical usage contexts in geotechnical engineering.
A categorization of the most basic DL architectures, using a mixture of both traditional and recently emerging DL architectures.
A detailed review of CNNs, LSTMs and GANs, which we have identified as the most commonly used methods in current geotechnical applications.
A discussion of transfer learning and its relevance for data-scarce geotechnical problems, with potential to improve model performance.
A collection of publicly available (i.e. open-source) geotechnical data sets with an overview of relevance and their potential use in DL-based studies.
The aim of clearly articulating each technique and providing practical resources is to support the goal of reducing the entry barriers for geotechnical researchers who may be early in their careers and interested in adopting DL approaches in their research. This is foundational material for future research that will investigate practical applications of these methods in geotechnical engineering. These future studies will help inform practical uses of DL and explore how adaptable these techniques are to address real-world geotechnical problems.
The first author appreciates the Civil Engineering Department of Lakehead University for providing an appropriate workspace and unrestricted access to research resources during his study visit in Canada, which significantly contributed to the preparation of this paper. Additionally, the authors acknowledge the assistance of chat.openai.com in proofreading the manuscript, ensuring clarity and coherence throughout the paper.

