Assumptions, strengths and limitations of major ML models in accounting
| Model family | Core assumptions | Primary strengths | Key limitations |
|---|---|---|---|
| Random forests (RF) |
|
|
|
| Gradient boosting machines (GBM, XGBoost, LightGBM, CatBoost) |
|
|
|
| AdaBoost |
|
|
|
| SVMs |
|
|
|
| Deep learning (NNs, CNNs, LSTMs, Transformers) |
|
|
|
| Unsupervised models (Clustering, Autoencoders, Anomaly Detectors) |
|
|
|
| Reinforcement Learning (RL) |
|
|
|
| Generative AI/LLMs |
|
|
|
| Model family | Core assumptions | Primary strengths | Key limitations |
|---|---|---|---|
| Random forests ( | Predictive structure dispersed across many weak, heterogeneous signals Pervasive nonlinear interactions and threshold effects Variance reduction through aggregation improves generalisation | Strong performance on noisy, tabular accounting data (financial ratios, governance indicators, journal-entry attributes) Models complex interactions without manual specification Robust to multicollinearity and outliers | Opaque ensemble structure complicating interpretability Difficult to justify in audit documentation without post-hoc tools (such as feature importance and partial dependency plots) Less effective for highly imbalanced label settings unless tuned carefully |
| Gradient boosting machines ( | Residual errors contain systematic information Complex relationships can be approximated through sequential error correction Combines many simple models to build a highly flexible, nonlinear prediction function | State-of-the-art performance in distress, fraud, misstatement, and going-concern prediction Handles high-dimensional feature spaces and subtle interaction structures Highly tuneable for different accounting tasks | Sensitive to label noise – problematic for fraud, misstatement, and Risk of overfitting without careful regularisation Interpretability limited because it relies on many small trees; explanations rely on approximations Requires explicit handling of class imbalance (weighting, resampling or focal loss) for rare-event accounting outcomes such as fraud and bankruptcy |
| AdaBoost | Misclassified observations are especially informative Outcome labels are sufficiently reliable for reweighting Decision boundaries can be sharpened by focusing on “hard” cases | Effective when labels are observable and accurate (e.g. internal control weaknesses, restatements) Performs well on structured, moderately complex tabular datasets where weak learners can reliably extract signals | Poor compatibility with noisy or strategically distorted labels Amplifies misclassifications in fraud and Less robust than |
Data are separable (linearly or via kernel transformation) Optimal classification boundary maximises margin | Good performance on smaller, structured datasets Effective in high-dimensional spaces when kernels are well specified | Difficult to scale to large transactional datasets Kernel choice strongly influences performance Limited interpretability and compatibility with audit documentation | |
| Deep learning ( | Predictive relationships are hierarchical, distributed, and nonlinear Representations can be learned directly from raw or minimally processed input Large datasets needed for generalisation | State-of-the-art for textual and narrative accounting data (MD&A, earnings calls, Captures tone, sentiment, obfuscation, and semantic context Flexible across sequential and unstructured data types | High opacity; internal logic difficult to document or defend to regulators Data-hungry and prone to overfitting in typical accounting sample sizes Risk of hallucination and instability in generative tasks |
| Unsupervised models (Clustering, Autoencoders, Anomaly Detectors) | Normal behaviour can be learned from majority patterns Distance/density metrics correspond to economic similarity Latent structure reflects meaningful categories or behaviours | Can detect anomalous journal entries, unusual transactions, and client-risk segmentation Do not require labelled outcomes – useful where fraud or misstatements are rarely observed | Discovered patterns may not correspond to accounting constructs High false-alarm risk; anomalies may reflect benign events Interpretation requires substantial professional judgement |
| Reinforcement Learning ( | Decisions shape future states and rewards Optimal policies can be learned through repeated interaction Organisational/compliance objectives are representable in reward functions | Conceptually promising for adaptive forecasting, audit effort allocation, and internal control optimisation Models complex sequential decision processes | Limited real-world adoption due to data scarcity and regulatory constraints Reward functions difficult to specify in normative accounting contexts Trial-and-error learning inappropriate for high-stakes environments |
| Generative AI/LLMs | Language encodes latent semantic and institutional structure Attention-based contextual embeddings capture meaning Scale produces emergent capabilities | Good at analysing narrative disclosures, contracts, and Useful for summarising, extracting terms, drafting workpapers Transforms linguistic information into quantifiable signals | Susceptible to hallucination and unverifiable reasoning Limited audit admissibility due to lack of transparent reasoning chains Requires strict governance for privacy, accuracy and accountability |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.