Table 1

Assumptions, strengths and limitations of major ML models in accounting

Model familyCore assumptionsPrimary strengthsKey limitations
Random forests (RF)
  • Predictive structure dispersed across many weak, heterogeneous signals

  • Pervasive nonlinear interactions and threshold effects

  • Variance reduction through aggregation improves generalisation

  • Strong performance on noisy, tabular accounting data (financial ratios, governance indicators, journal-entry attributes)

  • Models complex interactions without manual specification

  • Robust to multicollinearity and outliers

  • Opaque ensemble structure complicating interpretability

  • Difficult to justify in audit documentation without post-hoc tools (such as feature importance and partial dependency plots)

  • Less effective for highly imbalanced label settings unless tuned carefully

Gradient boosting machines (GBM, XGBoost, LightGBM, CatBoost)
  • Residual errors contain systematic information

  • Complex relationships can be approximated through sequential error correction

  • Combines many simple models to build a highly flexible, nonlinear prediction function

  • State-of-the-art performance in distress, fraud, misstatement, and going-concern prediction

  • Handles high-dimensional feature spaces and subtle interaction structures

  • Highly tuneable for different accounting tasks

  • Sensitive to label noise – problematic for fraud, misstatement, and GCO datasets, i.e. if the label is wrong, the residual is wrong – and the model learns from the error itself

  • Risk of overfitting without careful regularisation

  • Interpretability limited because it relies on many small trees; explanations rely on approximations

  • Requires explicit handling of class imbalance (weighting, resampling or focal loss) for rare-event accounting outcomes such as fraud and bankruptcy

AdaBoost
  • Misclassified observations are especially informative

  • Outcome labels are sufficiently reliable for reweighting

  • Decision boundaries can be sharpened by focusing on “hard” cases

  • Effective when labels are observable and accurate (e.g. internal control weaknesses, restatements)

  • Performs well on structured, moderately complex tabular datasets where weak learners can reliably extract signals

  • Poor compatibility with noisy or strategically distorted labels

  • Amplifies misclassifications in fraud and GCO contexts

  • Less robust than GBM or RF for high-stakes predictions

SVMs
  • Data are separable (linearly or via kernel transformation)

  • Optimal classification boundary maximises margin

  • Good performance on smaller, structured datasets

  • Effective in high-dimensional spaces when kernels are well ⁠specified

  • Difficult to scale to large transactional datasets

  • Kernel choice strongly influences performance

  • Limited interpretability and compatibility with audit documentation

Deep learning (NNs, CNNs, LSTMs, Transformers)
  • Predictive relationships are hierarchical, distributed, and nonlinear

  • Representations can be learned directly from raw or minimally processed input

  • Large datasets needed for generalisation

  • State-of-the-art for textual and narrative accounting data (MD&A, earnings calls, ESG disclosures)

  • Captures tone, sentiment, obfuscation, and semantic context

  • Flexible across sequential and unstructured data types

  • High opacity; internal logic difficult to document or defend to regulators

  • Data-hungry and prone to overfitting in typical accounting sample sizes

  • Risk of hallucination and instability in generative tasks

Unsupervised models (Clustering, Autoencoders, Anomaly Detectors)
  • Normal behaviour can be learned from majority patterns

  • Distance/density metrics correspond to economic similarity

  • Latent structure reflects meaningful categories or behaviours

  • Can detect anomalous journal entries, unusual transactions, and client-risk segmentation

  • Do not require labelled outcomes – useful where fraud or misstatements are rarely observed

  • Discovered patterns may not correspond to accounting constructs

  • High false-alarm risk; anomalies may reflect benign events

  • Interpretation requires substantial professional judgement

Reinforcement Learning (RL)
  • Decisions shape future states and rewards

  • Optimal policies can be learned through repeated interaction

  • Organisational/compliance objectives are representable in reward functions

  • Conceptually promising for adaptive forecasting, audit effort allocation, and internal control optimisation

  • Models complex sequential decision processes

  • Limited real-world adoption due to data scarcity and regulatory constraints

  • Reward functions difficult to specify in normative accounting contexts

  • Trial-and-error learning inappropriate for high-stakes environments

Generative AI/LLMs
  • Language encodes latent semantic and institutional structure

  • Attention-based contextual embeddings capture meaning

  • Scale produces emergent capabilities

  • Good at analysing narrative disclosures, contracts, and ESG reports

  • Useful for summarising, extracting terms, drafting workpapers

  • Transforms linguistic information into quantifiable signals

  • Susceptible to hallucination and unverifiable reasoning

  • Limited audit admissibility due to lack of transparent reasoning chains

  • Requires strict governance for privacy, accuracy and accountability

or Create an Account

Close subscription notice
Close access options