Article navigation
Purpose

The failure prediction models are attractive part of corporate finance for academics and practitioners alike. Traditionally, statistical methods have been used, but gradually machine learning (ML) models penetrate this field. Much of the studies using ML methods rely solely on financial ratios. The purpose of this paper is to compare the ML methods with traditional statistical models for different groups of explanatory variables and evaluate variable importance across all models.

Design/methodology/approach

The authors use data on Slovak small- and medium-sized enterprises (SMEs) in the pre-covid-19 period with 587,993 observations for 168,259 unique companies. Methods of logistic regression, bagging and boosting tree-based models are used to predict companies’ bankruptcy.

Findings

The results confirm that a broader set of variables improves performance of models and that, on average, well-tuned ML models are better than statistical ones. The difference between the two depends on the group of variables used for prediction. When comparing two ML models, the authors found that eXtreme Gradient Boosting becomes better than random forest when the group of features is expanded. The authors found that similar key variables have the greatest importance in all models.

Originality/value

This research analyses performance of traditional and ML methods in failure prediction based on different groups of features (using variety of financial and non-financial variables/features). At the same time, variable importance of multiple models is evaluated and compared using several techniques.

Licensed re-use rights only
You do not currently have access to this content.
Don't already have an account? Register

Purchased this content as a guest? Enter your email address to restore access.

Pay-Per-View Access
$39.00
Rental

or Create an Account

Close subscription notice
Close access options