The DNN architecture used in this study
| Hyperparameter | Applied | Justification |
|---|---|---|
| Normalisation method | StandardScaler | Ensures all features have comparable scales, improving convergence and numerical stability |
| Hidden layers | 3 fully connected layers (128, 64, 32 neurons) | Provides sufficient capacity to model nonlinear temporal and regional price dependencies without excessive complexity |
| Loss function | MSE | Appropriate for continuous-valued regression and forecasting tasks |
| Optimiser | Adam (learning rate = 0.001) | Adaptive optimiser, robust for non-stationary and noisy time series |
| Batch size | 4 | Small batch sizes improve generalisation for small-to-medium datasets |
| Epochs | 100 | Allows convergence while monitoring for overfitting, especially with small datasets |
| Input activation | Linear | Preserves scaled and PCA-transformed features without distortion |
| Hidden activation | ReLU | Efficient, avoids vanishing gradients, captures nonlinear dependencies |
| Output activation | Linear | Suitable for continuous-valued forecasting |
| Dropout rate | 0.2 | Prevents overfitting by randomly deactivating 20% of neurons during training |
| Batch Normalisation | Enabled | Stabilises learning, accelerates convergence, reduces internal covariate shift |
| Early stopping | Enabled | Prevents overfitting and avoids unnecessary computation once validation loss plateaus |
| Data partitioning method | Holdout | Partition the dataset into 70% for training, 10% for validation, and 20% for testing |
| Hyperparameter | Applied | Justification |
|---|---|---|
| Normalisation method | StandardScaler | Ensures all features have comparable scales, improving convergence and numerical stability |
| Hidden layers | 3 fully connected layers (128, 64, 32 neurons) | Provides sufficient capacity to model nonlinear temporal and regional price dependencies without excessive complexity |
| Loss function | MSE | Appropriate for continuous-valued regression and forecasting tasks |
| Optimiser | Adam (learning rate = 0.001) | Adaptive optimiser, robust for non-stationary and noisy time series |
| Batch size | 4 | Small batch sizes improve generalisation for small-to-medium datasets |
| Epochs | 100 | Allows convergence while monitoring for overfitting, especially with small datasets |
| Input activation | Linear | Preserves scaled and PCA-transformed features without distortion |
| Hidden activation | Efficient, avoids vanishing gradients, captures nonlinear dependencies | |
| Output activation | Linear | Suitable for continuous-valued forecasting |
| Dropout rate | 0.2 | Prevents overfitting by randomly deactivating 20% of neurons during training |
| Batch Normalisation | Enabled | Stabilises learning, accelerates convergence, reduces internal covariate shift |
| Early stopping | Enabled | Prevents overfitting and avoids unnecessary computation once validation loss plateaus |
| Data partitioning method | Holdout | Partition the dataset into 70% for training, 10% for validation, and 20% for testing |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.