The purpose of this study is twofold. Firstly, this study outlines the creation of an ML “bolt on” transaction monitoring risk scoring model. This ML model uses predictive classification to score the likelihood of an alert being a false or true positive based on historical data. Secondly, this study explores the results of implementing the model in a real-world environment. This is valuable, as it explores the potential to reduce the workload on human analysts responsible for reviewing alerts, by “hibernating” low-risk alerts and “auto-escalating” high-risk alerts.
Machine learning is used to create a “bolt on” model that sits on top of existing rule-based transaction monitoring systems to improve their effectiveness and efficiency. This was achieved by developing a model to mimic the analysts alert review process with the aim of improving efficiency, reducing investigation timeliness and ultimately reducing time-to-SAR duration. The model did so by scoring anti-money laundering (AML) alerts to help conduct three actions: automatic hibernation of low-risk alerts; review of medium-risk alerts by the level 1 team; and auto-escalation of high-risk alerts to the level 2 team for detailed investigation.
The model was successful in consistently identifying low-risk alerts correctly, significantly reducing the volume of false positives that needed manual review. Prior to implementation, 100% of alerts had to be reviewed, by the end of the examined period, 19% of alerts were hibernated and 18% of alerts were auto-escalated. Time taken to investigate and submit SARs was reduced by 61%.
The practical implications of the study lie in its contributions to helping banking institutions effectively keep pace with the increasing volume of transactions, in a cost effective and efficient manner, which has a positive effect on their reputation through enhancing their AML response.
The model outlined greatly enhances the timeliness of investigations, which is crucial for money laundering prevention and enables timely responses by institutions and law enforcement agencies to make better use of the investigation timeframe after an offence is identified, enabling them to more rapidly address threats, freeze assets when necessary, and secure vital evidence before it is lost.
This study provides a blueprint for other financial institutions to consider similar approaches to AML. This is significant because previous studies of machine learning applications in the domain of AML have not explored this type of use before, particularly in the context of machine learning ensembled techniques. It also addressed several cited limitations of previous research by increasing the depth of information in the datasets used and providing a formal evaluation. Importantly, the blueprint and evaluation metrics used in this study ensure its reproducibility, addressing another previously identified academic issue.
Introduction
Money laundering is defined by the 1998 UN Vienna Convention (Article 3.1) as “the conversion or transfer of property knowing that such property is derived from any offense(s), for the purpose of concealing or disguising the illicit origin of the property”, or put another way, money laundering is the act of making dirty money look clean.
To detect such activity, banks use transaction monitoring (TM) systems to analyse patterns of transactions within their institutions. However, TM systems are renowned for creating high volumes of false positives, which require large numbers of human resources in the form of financial analysts to review. These ever-increasing volumes are forcing banks to innovate to meet demand.
This article outlines and examines the performance of a “bolt-on” alert scoring model used by a major bank. For commercial privacy, both the country and institute involved remain anonymous throughout this study. Using machine learning (ML), the model automatically ranks suspicious activity as low, medium or high-risk enabling large volumes to be managed more efficiently through auto-hibernation or escalation for priority attention. This article examines the design and development in detail and provides a robust evaluation of performance and a blueprint for the financial industry to consider implementation of such models more widely.
Background
According to the United Nations Office on Drugs and Crime, a staggering $1.6tn is laundered each year (Global Coalition to Fight Financial Crime, 2021). Money laundering has significant negative macroeconomic and social consequences. The potential consequences of money laundering include increased exposure to organized crime; undermining the legitimate private sector; weakening of financial organizations; dampening of foreign investments; and economic distortion and instability (Chaikin and Sharman, 2009).
While money laundering itself is a crime in most areas, it is preceded by other illegal activity that is the foundation (aka predicate offense). For example, the illicit act of tax evasion is a predicate offense, and the illegal proceeds generated would be laundered. Predicate offenses may vary from country to country and are usually codified in a country’s criminal code. For instance, in the USA, these offenses were initially created by the Bank Secrecy Act of 1970 and have been expanded by the 2001 USA’s Patriot Act (Chitimira and Munedzi, 2023). In 2021, the European Union (EU) issued the Sixth EU Anti-Money Laundering Directive which defines and standardizes 22 predicate offenses including narcotrafficking, tax evasion, murder, grievous bodily harm, corruption, fraud, smuggling, human trafficking, illegal wildlife trafficking and forgery, among others (EU Directive 2015/849, 2021).
Placement, layering and integration
The act of money laundering is traditionally recognized as being conducted through a three-stage process consisting of placement, layering and integration (Cassella, 2018). In stage one, placement, funds derived from criminal activity are introduced into the financial system (Cassella, 2018). Often, this is accomplished by placing the funds into circulation through financial institutions, money service bureaus, or cash-based businesses (Cassella, 2018). Examples of placement include mixing illegitimate and legitimate funds together, such as placing the cash from illegal drug sales into a restaurant; purchasing significant numbers of stored value cards with cash; dividing cash into lesser amounts and depositing it into numerous bank accounts (Cassella, 2018).
In stage two, layering, the funds are repeatedly moved through the financial system to obfuscate the source of the funds, which makes it more difficult to identify the origin (Cassella, 2018). Examples of layering include moving funds from one financial institution to another; moving funds within accounts at the same institution; placing funds in stocks, bonds or insurance products; reselling high-value goods, prepaid store cards or other stored value products (Cassella, 2018).
The final stage, integration, is used to re-integrate the newly “cleaned” funds so that they can be used for normal purposes (Cassella, 2018). For example, the launderer might choose to invest the funds in real estate, financial ventures or luxury assets.
Financial Action Task Force
Formed in 1989, the Financial Action Task Force (FATF) is an intergovernmental body created by the group of seven industrialized nations, to set standards and foster international action against money laundering. In total, more than 200 countries and authorities have committed to implement the FATF’s standards as part of a coordinated global response to preventing organized crime and corruption. As a result, countries worldwide have established frameworks that obligate financial institutions to identify and report suspicious financial activity (Kern, 2000).
Money laundering typologies
Since its inception, the FATF has undertaken the study of money laundering typologies (i.e. methods and trends) as a key component of its work. In doing so, they have identified that common money laundering typologies include rapid movement of funds, structuring and use of dormant accounts (Gallo and Juckes, 2005). Rapid Movement of Funds involves the swift transfer of funds between various accounts or financial institutions, often without a clear business or legitimate reason. This behaviour can indicate attempts to obscure the origin of funds. Structuring involves the use of deposits or withdrawals just below regulatory reporting requirements to evade detection, and nay indicate suspicious activity (Gallo and Juckes, 2005). Furthermore, dormant accounts suddenly being used for elevated levels of activity, or multiple accounts used to move funds in a potential effort to obscure the source of money, are both common money laundering typologies (Ai and Tang, 2011; and Lai, 2010). When institutions suspect the presence of money laundering, often by identifying the outlined customer behaviour, they are required to submit a suspicious activity report (SAR). Failure to do so may make them liable to criminal or civil penalties (Pavlidis, 2023).
Reporting suspicious activity
SARs are a tool provided by the Bank Secrecy Act of 1970. Originally called a criminal referral form, the SAR has become the standard form to report suspicious activity. The purpose of SARs is to initiate the pro-active investigation of potentially illegal activity (Ping, 2005). SARs are the key vehicle that financial institutions use to highlight suspected money laundering to the authorities and Financial Intelligence Units (FIUs) (Ping, 2005; Stalcup, 2015). FIUs analyse SARs, and if they suspect criminal conduct, the information is then shared with the appropriate law enforcement agencies for further investigation (Pavlidis, 2023; and Ping, 2005).
Automated transaction monitoring systems
Spotting suspected money laundering is difficult as large institutions can process millions of transactions per day. Identification is heavily reliant on internal practices that are set up to identify potential indicators that suspicious activity is occurring (Hopkins and Shelton, 2019). The process of generating SARs starts with the detection of unusual activity by TM programmes, systems and processes designed to highlight behaviour that meets predefined typologies (Repin et al, 2017; Sapozhnikova et al., 2017, 2018). Large financial institutions deploy automated TM systems to detect money laundering (Chau and van Dijck Nemcsik, 2020). These TM systems flag transactions as “alerts” if they meet the criteria of the typologies/rules/scenarios configured (Harvey, 2020).
Alert management
The task of reviewing alerts falls to employees trained in analysis and examination of transactions for signs of illicit activity. The number of analysts required to review alerts can be considerable and large financial institutions can employee hundreds of analysts. These financial analysts are employed within alert management teams, who are typically structured into three sub-teams known as level 1, level 2 and level 3. The level 1 review is sometimes referred to as an “alert Investigation”. To conduct a level 1 review, an analyst may utilize various sources of internal and external data, including customer records, payment data, public domain searches, government and third-party databases. If the analyst can determine that the alert poses no concern, then they can close the alert. If not, the alert would be escalated to level 2, which results in the creation of a “case”.
This level 2 review, referred to as a “case investigation” is where an analyst conducts further analysis of the activity. These reviews typically require the procurement of additional data from the customer, other internal functions, or use of tools that level 1 analysts may not have access. If a level 2 analyst is unable to obtain the required information to close the alert, then the analyst will recommend escalation to level 3 for the filing of a SAR.
As part of the level 3 review, a SAR is prepared and submitted to the FIU. As an example of scale, in the UK, over 460,000 SARs are submitted every year. Collectively, these reports contribute a great deal of information towards the anti-money laundering (AML) efforts of financial authorities. A typical SAR is less than 10 pages but a SAR for a complex transaction could be as long as 50 pages.
Optimizing transaction monitoring systems
A limitation of TM systems is that they may not always efficiently detect money laundering, which leads to a high volume of false-positive alerts. A typical TM system may have an alert-to-SAR ratio of just 5%, and to address this issue, it is necessary to refine their performance (Ratanachu-ek, 2021). Two techniques traditionally used to achieve this are segmentation and tuning (Ratanachu-ek, 2021).
Segmentation involves dividing customers into homogeneous groups based on similar transaction behaviours or risk profiles, allowing for threshold settings for distinct groups (Ratanachu-ek, 2021). This ensures banks have homogenous groups who possess similar behaviour, products, and/or services, and therefore, unusual activity is easier to identify. This method has inherent risks if not done correctly. For example, if all banking customers were placed in the same group, then this could result in comparisons between ordinary customers, students, high-net-worth individuals or even businesses. Clearly, these groups have little in common in terms of their financial transactional behaviour and if compared against one another may negatively affect the rates of false positives.
Tuning involves regularly analysing transactional data to set and adjust parameters, which minimizes false positives, while also ensuring significant alerts are not overlooked (Ratanachu-ek, 2021). For instance, what constitutes a “large transaction” for one customer, may be modest for another, necessitating different monitoring parameters for both. Parameters can be set for a range of features and examples include “transaction size”, “maximum monthly balance” and “number of cash deposits”. Setting the parameter values for certain scenarios can be quite simple, as some are defined by law. For instance, the threshold value for what constitutes a large transaction is usually outlined in localized legislation. However, other parameters must be set using local expertise or derived from a statistical approach that focuses on behaviours calculated to identify anomalies.
Risk scoring of alerts
To help alert management teams assess risk, TM systems often assign risk scores to alerts based on a broad range of factors that are indicative of suspicious activity (Simpson, 2018). Risk scoring allows the institute to identify high-risk alerts quickly (Simpson, 2018). In recent years, the efficiency of TM systems has been significantly enhanced by technologies such as artificial intelligence (AI) (Gupta et al., 2023; Landge et al., 2024; Zhang and Chen, 2024), including ML (Gerlings and Constantiou, 2022; Oztas et al., 2022), and has enabled large organizations to deploy complex data mining techniques to improve the accuracy of detection.
Machine learning research on transaction monitoring and risk assessment
The ease of digital banking has caused the volume of transactions monitored by financial institutions to dramatically increase in recent years (Oztas et al., 2022). The availability of digital data has also significantly enhanced the TM process as digital data provides many opportunities for identifying patterns and trends in activities that humans may overlook (Lokanan and Maddhesia, 2023). Studies on ML techniques applied in this domain have historically focused on detecting unidentified anomalies using historical AML cases (Lv et al., 2008), graph analysis (Alshantti and Rasheed, 2021; Desrousseaux et al., 2021; Rui and Wunsch, 2005), anomaly detection (Gao, 2009; Liu et al., 2008; Muhammed Shokry et al., 2020), behavioural analysis (Ketenci et al., 2021; Jullum et al., 2020) and risk rating (Larik and Haider, 2011). As such, the effectiveness of TM systems has been enhanced in recent years, especially with the use of AI and ML (Gupta et al., 2023; Landge et al., 2024; Zhang and Chen, 2024; Gerlings and Constantiou, 2022; Oztas et al., 2022).
Despite these studies, a review of ML techniques that improve TM has identified significant gaps in literature (Oztas et al., 2022). Specifically, there is a lack of research in the use of ensembled learning, reinforcement and deep learning methods; studies lacking information or reproducibility; minimal information on data sets used; and limited evaluation (Oztas et al., 2022). The review concluded that solutions to detect money laundering remain an ongoing research area, with insufficient literature available regarding ML methods for transactional monitoring, especially using real data sets (Oztas et al., 2022).
The current study
The purpose of this study is twofold. Firstly, we outline the creation of an ML “bolt on” TM risk scoring model. This ML model uses predictive classification to score the likelihood of an alert being a false or true positive based on historical data (Ratanachu-ek, 2021). Secondly, we explore the results of implementing the model in a real-world environment. This is valuable as it explores the potential to reduce workload on human analysts responsible for reviewing alerts, by “auto-hibernating” low-risk alerts and “auto-escalating” high-risk alerts.
Exploring these two objectives is important for several reasons. First, it may provide a blueprint of how to develop similar models. If effective, then these may help financial institutions increase capacity to detect money laundering. Second, it could also reduce the time taken to investigate suspicious activity, enhancing the institutions AML response. Finally, the use of this method could refine TM systems and extend this under researched area by outlining one of the first examinations of the alert scoring method for optimizing a TM risk assessment process.
Development of the risk scoring model for managing money laundering
“Bolt on” technologies that sit on top of rule-based systems have been cited as a method to improve efficiency of TM systems by using ML to risk-score and auto-hibernate or escalate alerts (Oztas et al., 2024). We achieved this by developing a model to mimic the analysts alert review process with the aim of improving efficiency, reducing investigation timeliness and improving the time-to-SAR duration. The model did so by scoring AML alerts to help conduct three actions:
automatic hibernation of low-risk alerts;
review of medium-risk alerts by the level 1 team; and
automatic escalation of high-risk alerts to the level 2 team for detailed investigation.
Building and hyper-tuning the machine learning model
To score and provide a rank-ordered list of alerts, the bank used an ensemble ML technique called boosting which constructs, evaluates and tests the model to create a stronger and more accurate prediction model. The model used the LightGBM gradient boosting framework for its fast-training speed, low memory usage, accuracy and ability to handle large data sets. Comparative studies have demonstrated that LightGBM performs competitively with other ensemble learning algorithms like XGBoost and Random Forest when used for classification tasks (Li and Chen, 2020). LightGBM was also selected due to its ability to impose monotonic constraints on variables. This functionality to iteratively over-penalize erroneous decisions at each tree-building stage contributed to improved model performance.
The data used to build and train the model was collected from three years of TM alerts, compliance flags, SAR data, customer demographic details and transactional data. The model was trained on the relationship between analysts’ historical decisions and the associated customer, account and transactional behaviour. This was achieved by conducting iterative experimentation and model training on the historical data.
While all alerts are considered suspicious, those leading to SARs send a much stronger signal. To leverage this, the model oversamples alerts that produced SARs by three times, allowing the estimator to emphasize this signal during the learning process. The goal of doing so was to ensure that the estimator iteratively over-penalized erroneous false negative decisions for SARs at each tree build, which reduced false negatives.
A predictive binary classification model was developed to risk score alerts. Initially, the data contained approximately 2,000 predictive variables (features), which were assessed based on their contribution towards performance. These were later reduced to improve and hyper tune the model. The hyperparameter grid related to this process can be found in Appendix Table A1.
Hyperparameter grid
| Hyperparameter | Description | Search space | Model |
|---|---|---|---|
| num_iterations | Number of boosting iterations | Fixed to 100 by the developers to Reduce down the search space to Facilitate efficient optimization Learning_rate hyperparameter is Tuned as a proxy for the effect of Lesser or higher number of trees | 100 |
| feature_fraction | LightGBM will randomly select A subset of features on each Iteration (tree), if feature fraction is smaller than 1.0. For example, if you set it to 0.8, LightGBM will select 80% Of features before training each tree • can be used to speed up training • can be used to deal with over-fitting | 0.3, 0.5, 0.7, 0.9 | 0.3 |
| max_depth* * Important hyperparameter | Limits the max depth of a tree This is used to deal with overfitting when #data is small. Tree Still grows leaf-wise | 2, 4, 8, 10 | 4 |
| learning_rate* * Important hyperparameter | A technique to slow down the Learning by weighting the Corrections by new trees at Every learning iteration | 0.001, 0.01, 0.05, 0.1 | 0.1 |
| min_gain_to_split | Minimal gain to perform | 0.3, 0.5, 0.7 | 0.3 |
| min_data_in_leaf | Minimal number of data in one Leaf. Can be used to deal with over-fitting | 5, 10, 15, 20 | 15 |
| lambda_l1 | L1 regularization | 0.1, 0.3 | 0.3 |
| scale_pos_weight | Weight of labels with positive class | 1.0, 1.5, 1.7 | 1.7 |
| Hyperparameter | Description | Search space | Model |
|---|---|---|---|
| num_iterations | Number of boosting iterations | Fixed to 100 by the developers to | 100 |
| feature_fraction | LightGBM will randomly select | 0.3, 0.5, 0.7, 0.9 | 0.3 |
| max_depth* | Limits the max depth of a tree | 2, 4, 8, 10 | 4 |
| learning_rate* | A technique to slow down the | 0.001, 0.01, 0.05, 0.1 | 0.1 |
| min_gain_to_split | Minimal gain to perform | 0.3, 0.5, 0.7 | 0.3 |
| min_data_in_leaf | Minimal number of data in one | 5, 10, 15, 20 | 15 |
| lambda_l1 | L1 regularization | 0.1, 0.3 | 0.3 |
| scale_pos_weight | Weight of labels with positive | 1.0, 1.5, 1.7 | 1.7 |
Source(s): Authors’ own creation/work
Feature selection
The feature selection methodology involved multiple sequential steps. While the initial model performed satisfactorily, further simplification was pursued to reduce complexity. Feature importance scores generated by LightGBM help identify the most influential features in integrated classification models (Huang and Wang, 2023). Using LightGBM feature importance method, the lowest 10% of features were iteratively removed until the model performance was compromised. This enabled us to reduce the initial set of 2,000 features to 356 that could be assessed based on contribution towards performance.
At this point it became evident that most feature variables represented customer transactional behaviour calculated over different lookback windows, leading to a set of highly multicollinear variables. Once this point was reached, we incorporated insights from the banks compliance experts to further narrow down the feature variables retaining only those with the highest correlation to the target variable and reducing the final list to 91. A summary description is provided in Table 1.
High-level feature variables
| Types of variables | Rationale | Examples |
|---|---|---|
| Customer features | Unique customer elements could provide specific risk indicators associated with client’s demographics | Sector class, geographic location, number of years of client relationship, employment type, expected turnover, etc |
| Product use | Features related to various type of transactions identifies different risks | Cash intensive, high-risk countries, rapid movements, round amounts, channels, recency, risky merchant |
| Compliance features | Customers who are already part of any other investigation or watchlist or STR can be escalated to L2 directly | Previous STR, corr bank queries, watchlist, CIU or EDD concerns, payment declines |
| Relationships | Any unusual relationship based on counterparties or related parties | Number of counterparties, risky counterparty (STR), STR on related party |
| Existing AML rules | Indicators based on the past alerts and each detection scenario | DS flags, number of previous alerts |
| Types of variables | Rationale | Examples |
|---|---|---|
| Customer features | Unique customer elements could provide specific risk indicators associated with client’s demographics | Sector class, geographic location, number of years of client relationship, employment type, expected turnover, etc |
| Product use | Features related to various type of transactions identifies different risks | Cash intensive, high-risk countries, rapid movements, round amounts, channels, recency, risky merchant |
| Compliance features | Customers who are already part of any other investigation or watchlist or STR can be escalated to L2 directly | Previous STR, corr bank queries, watchlist, CIU or EDD concerns, payment declines |
| Relationships | Any unusual relationship based on counterparties or related parties | Number of counterparties, risky counterparty (STR), STR on related party |
| Existing AML rules | Indicators based on the past alerts and each detection scenario | DS flags, number of previous alerts |
Source(s): Authors’ own creation/work
Machine learning performance and validation
Table 2 provides the training and out-of-time (OOT) testing sample periods, sample counts and positive class rates for the model. The model was trained using 16 months of data from January 2021 to April 2022 and validated using 12 months of data from May 2022 to April 2023. The training sample size was 59,685, while the validation sample size was 105,589. Positive class rates were consistent across training and validation periods, with the model showing 34% positive cases and 30% in validation. This consistency indicates stable data distribution over time. In addition, the training sample included 3,281 SARs, while the validation sample included 3,190 SARs. This substantial number of SARs in the training and validation datasets ensured that the model could learn and validate effectively. Overall, the table highlights that the model has adequate sample sizes and maintains performance consistency across different time periods, ensuring robust training and evaluation.
Training and OOT testing sample periods, sample counts and positive class rates
| Factor | Training | Validation |
|---|---|---|
| Period | [Jan-21: Apr-22] | [May-22: Apr-23] |
| Sample size | 59,685 | 105,589 |
| Positive class rate | 0.34 | 0.30 |
| SAR count | 3,281 | 3,190 |
| Factor | Training | Validation |
|---|---|---|
| Period | [Jan-21: Apr-22] | [May-22: Apr-23] |
| Sample size | 59,685 | 105,589 |
| Positive class rate | 0.34 | 0.30 |
| SAR count | 3,281 | 3,190 |
Source(s): Authors’ own creation/work
We also used receiver operator characteristic (ROC) curves and the area under the curve (AUC) to evaluate the model performance. ROC curves help assess how well the model distinguishes between true and false positives, determine the optimal threshold for classification and compare performance of different models. LightGBM’s hyperparameters provided a mechanism for maintaining balance between model underfitting and overfitting. RandomizedSearchCV from scikit-learn was used to search the parameter grid.
The difference between training and test ROC AUC scores (diff_roc_auc) was selected as the objective function of the optimization analysis. We aimed to minimize diff_roc_auc to ensure consistent model performance across training and test data sets. For the model, the selected hyperparameter configuration resulted in a diff_roc_auc of approximately 0.03, with a mean training ROC AUC of around 0.86 and a mean test ROC AUC of around 0.83. These results, shown in Appendix Figure A1, indicate that the model maintained reliable performance with minimal overfitting, as evidenced by the minimal differences between training and test ROC AUC scores.
Model ranking performance
As the model must rank alerts, this was also calibrated. The ROC, and AUC are again selected as performance evaluation metrics. The ROC curve is used as it plots the true positive rate (TPR) against the false positive rate (FPR) at various threshold settings. The AUC provides a single-number summary of the curve, ranging from 0.0 to 1.0, where 0.5 represents random model performance and 1.0 represents perfect model performance. The model was back-tested over four OOT test periods. Each OOT sample contained three months of alert data generated from May 2022 to April 2023. The Model Classification report in Appendix Table A2 lists the model AUC values for all OOT tests, with ROC curve figures provided in Appendix Figure A2. Metrics for validating TM alerts include precision, TPR of level 2 cases and TPR and false negative rates (FNR) of SAR at top percentiles of prediction scores, such as the top 10, 20 and 30 percentiles.
Model classification report
| Top percentile | isl2_precision | isl2_tpr | sar_tpr | sar_fnr | l2_roc_auc | sar_roc_auc |
|---|---|---|---|---|---|---|
| 10 | 0.882558372 | 0.241146 | 0.654776 | 0.3452237 | 0.8248399 | 0.9139197 |
| 20 | 0.784563189 | 0.428719 | 0.833938 | 0.1660621 | 0.8248399 | 0.9139197 |
| 30 | 0.710218201 | 0.582131 | 0.91576 | 0.0842402 | 0.8248399 | 0.9139197 |
| 40 | 0.646002245 | 0.70599 | 0.961104 | 0.0388956 | 0.8248399 | 0.9139197 |
| 50 | 0.590907277 | 0.807219 | 0.984684 | 0.0153164 | 0.8248399 | 0.9139197 |
| 60 | 0.535366807 | 0.877614 | 0.992946 | 0.0070536 | 0.8248399 | 0.9139197 |
| 70 | 0.486400182 | 0.930233 | 0.996977 | 0.003023 | 0.8248399 | 0.9139197 |
| 80 | 0.443762552 | 0.969928 | 0.998992 | 0.0010077 | 0.8248399 | 0.9139197 |
| 90 | 0.404002661 | 0.993402 | 1 | 0 | 0.8248399 | 0.9139197 |
| Top percentile | isl2_precision | isl2_tpr | sar_tpr | sar_fnr | l2_roc_auc | sar_roc_auc |
|---|---|---|---|---|---|---|
| 10 | 0.882558372 | 0.241146 | 0.654776 | 0.3452237 | 0.8248399 | 0.9139197 |
| 20 | 0.784563189 | 0.428719 | 0.833938 | 0.1660621 | 0.8248399 | 0.9139197 |
| 30 | 0.710218201 | 0.582131 | 0.91576 | 0.0842402 | 0.8248399 | 0.9139197 |
| 40 | 0.646002245 | 0.70599 | 0.961104 | 0.0388956 | 0.8248399 | 0.9139197 |
| 50 | 0.590907277 | 0.807219 | 0.984684 | 0.0153164 | 0.8248399 | 0.9139197 |
| 60 | 0.535366807 | 0.877614 | 0.992946 | 0.0070536 | 0.8248399 | 0.9139197 |
| 70 | 0.486400182 | 0.930233 | 0.996977 | 0.003023 | 0.8248399 | 0.9139197 |
| 80 | 0.443762552 | 0.969928 | 0.998992 | 0.0010077 | 0.8248399 | 0.9139197 |
| 90 | 0.404002661 | 0.993402 | 1 | 0 | 0.8248399 | 0.9139197 |
Note(s): ROC AUC of L2 and SAR classes are 0.825 and 0.914, respectively
Source(s): Authors’ own creation/work
Model implementation
The first phase of the model implementation was conducted as a six-month project to ensure that the model could be tested using real-time customer data without any associated risk from false negatives and positives.
During the pilot, for low-risk alerts, analysts were instructed to expedite their investigations using a “light-touch” review process which lasted 5 min versus the typical process which is considerably longer. For high-risk alerts, investigators were instructed to skip the level 1 review entirely and immediately escalate to level 2.
Throughout the pilot, results were monitored to track alignment between the model’s risk scores and decisions made by human analysts. On conclusion of the pilot, and after establishing successful performance of the model, the full implementation was conducted, and new working practices were put into place. The light-touch review was replaced by auto-hibernation of low-risk alerts and the threshold was increased to 20% of alerts and subsequently further increased to 30%.
In addition, overrides were incorporated into the auto-decisioning process that meant certain alerts could not be hibernated. For instance, if a customer had previously been the subject of a SAR, then future alerts could not be hibernated. Likewise, if a customer had two low scoring alerts hibernated in the previous 90 days, then the third alert could not be hibernated.
To assess the results, we used 11 months of continuous data, covering both the pilot and implementation period to assess performance. To understand improvements in efficiency, we use the data to derive time saved in analyst reviews.
Findings
Results indicate that the implementation of the risk scoring model for managing money laundering detection has been successful. Tables 3 and 4 outline the descriptive results obtained for low- and high-risk alerts, respectively, and indicate that the model performed extremely well.
Performance on assessment of low-risk alerts
| Month | Total alerts | No. of low-risk alerts | % Low-risk alerts | No. closed without SAR | % Closed at without SAR | Missed SARS | FTE’s Saved |
|---|---|---|---|---|---|---|---|
| Apr-23 | 9,438 | 820 | 9 | 820 | 100 | 0 | 6.8 |
| May-23 | 5,487 | 412 | 8 | 412 | 100 | 0 | 3.4 |
| Jun-23 | 9,285 | 1,322 | 14 | 1,322 | 100 | 0 | 11.0 |
| Jul-23 | 10,156 | 1,394 | 14 | 1,394 | 100 | 0 | 11.6 |
| Aug-23 | 9,866 | 1,522 | 15 | 1,522 | 100 | 0 | 12.7 |
| Sep-23 | 11,274 | 1,093 | 10 | 1,093 | 100 | 0 | 9.1 |
| Oct-23 | 10,315 | 1,184 | 11 | 1,184 | 100 | 0 | 9.9 |
| Nov-23 | 12,123 | 1,241 | 10 | 1,241 | 100 | 0 | 10.3 |
| Dec-23 | 12,831 | 1,931 | 15 | 1,931 | 100 | 0 | 16.1 |
| Jan-24 | 12,179 | 1,965 | 16 | 1,965 | 100 | 0 | 16.4 |
| Feb-24 | 14,540 | 2,812 | 19 | 2,812 | 100 | 0 | 23.4 |
| Total | 117,494 | 15,696 | 13.4 | 15,696 | 100 | 0 | 130.8 |
| Month | Total alerts | No. of low-risk alerts | % Low-risk alerts | No. closed without SAR | % Closed at without SAR | Missed SARS | FTE’s |
|---|---|---|---|---|---|---|---|
| Apr-23 | 9,438 | 820 | 9 | 820 | 100 | 0 | 6.8 |
| May-23 | 5,487 | 412 | 8 | 412 | 100 | 0 | 3.4 |
| Jun-23 | 9,285 | 1,322 | 14 | 1,322 | 100 | 0 | 11.0 |
| Jul-23 | 10,156 | 1,394 | 14 | 1,394 | 100 | 0 | 11.6 |
| Aug-23 | 9,866 | 1,522 | 15 | 1,522 | 100 | 0 | 12.7 |
| Sep-23 | 11,274 | 1,093 | 10 | 1,093 | 100 | 0 | 9.1 |
| Oct-23 | 10,315 | 1,184 | 11 | 1,184 | 100 | 0 | 9.9 |
| Nov-23 | 12,123 | 1,241 | 10 | 1,241 | 100 | 0 | 10.3 |
| Dec-23 | 12,831 | 1,931 | 15 | 1,931 | 100 | 0 | 16.1 |
| Jan-24 | 12,179 | 1,965 | 16 | 1,965 | 100 | 0 | 16.4 |
| Feb-24 | 14,540 | 2,812 | 19 | 2,812 | 100 | 0 | 23.4 |
| Total | 117,494 | 15,696 | 13.4 | 15,696 | 100 | 0 | 130.8 |
Source(s): Authors’ own creation/work
Performance on assessment of high-risk alerts
| Month | Total alerts | High-risk alerts | % High-risk alerts | No. of SARs | SAR coverage | FTEs Saved |
|---|---|---|---|---|---|---|
| Apr-23 | 9,438 | 816 | 8.6 | 324 | 47 | 6.8 |
| May-23 | 5,487 | 376 | 6.9 | 142 | 39 | 3.1 |
| Jun-23 | 9,285 | 788 | 8.5 | 328 | 56 | 6.6 |
| Jul-23 | 10,156 | 1,590 | 15.7 | 472 | 58 | 13.3 |
| Aug-23 | 9,866 | 1,583 | 16.0 | 416 | 54 | 13.2 |
| Sep-23 | 11,274 | 1,946 | 17.3 | 425 | 55 | 16.2 |
| Oct-23 | 10,315 | 1,715 | 16.6 | 331 | 49 | 14.3 |
| Nov-23 | 12,123 | 2,127 | 17.5 | 471 | 55 | 17.7 |
| Dec-23 | 12,831 | 2,142 | 16.7 | 499 | 67 | 17.9 |
| Jan-24 | 12,179 | 2,211 | 18.2 | 497 | 66 | 18.4 |
| Feb-24 | 14,540 | 2,660 | 18.3 | 558 | 67 | 22.2 |
| Total | 117,494 | 17,954 | 15.3 | 4,463 | 57 | 149.6 |
| Month | Total alerts | High-risk alerts | % High-risk alerts | No. of SARs | SAR coverage | FTEs |
|---|---|---|---|---|---|---|
| Apr-23 | 9,438 | 816 | 8.6 | 324 | 47 | 6.8 |
| May-23 | 5,487 | 376 | 6.9 | 142 | 39 | 3.1 |
| Jun-23 | 9,285 | 788 | 8.5 | 328 | 56 | 6.6 |
| Jul-23 | 10,156 | 1,590 | 15.7 | 472 | 58 | 13.3 |
| Aug-23 | 9,866 | 1,583 | 16.0 | 416 | 54 | 13.2 |
| Sep-23 | 11,274 | 1,946 | 17.3 | 425 | 55 | 16.2 |
| Oct-23 | 10,315 | 1,715 | 16.6 | 331 | 49 | 14.3 |
| Nov-23 | 12,123 | 2,127 | 17.5 | 471 | 55 | 17.7 |
| Dec-23 | 12,831 | 2,142 | 16.7 | 499 | 67 | 17.9 |
| Jan-24 | 12,179 | 2,211 | 18.2 | 497 | 66 | 18.4 |
| Feb-24 | 14,540 | 2,660 | 18.3 | 558 | 67 | 22.2 |
| Total | 117,494 | 17,954 | 15.3 | 4,463 | 57 | 149.6 |
Source(s): Authors’ own creation/work
Low-risk model
With respect to low-risk alerts, we can see in Table 3, on average, the model categorized approximately 13% of alerts as low-risk (n = 15,696). During the first 2 months the number was approximately 9%, but this gradually increased to 19% in subsequent months. There were no SARs missed during the six-month pilot during the light-touch review process. After five months the light-touch review process was replaced with automatic hibernation of “low-risk” alerts and the model continued to perform extremely well, not missing any SARs.
High-risk model
With respect to high-risk alerts, we can see that on average the model categorized approximately 15% of alerts as high-risk (n = 17,954). During the first 3 months the percentage was approximately 7%–9%, and this was gradually increased in subsequent months to over 18%. As the percentage of alerts tagged as high risk was increased, the percentage of total SARs covered also increased.
In this context, we see the model performed excellently throughout the entire period, with an average of 57% of the total SARs reported during this period being covered by top 15% of alerts.
As discussed, the aim of the model was to enhance the efficiency of investigations into suspicious activity. The best way to measure this is the timeliness of investigations, as delays result in missed opportunities for investigators to potentially freeze assets and funds held by the bank, and secure vital evidence. Figure 1 illustrates the lead time between generating an alert and submitting a SAR has dramatically reduced by 61%.
Time savings
To help understand if the implementation of the model improved the efficiency of the bank, we equated the time saved into human resources. This is achieved by calculating the number of alerts processed by employees. For example, without the model, an employee usually processed approximately 120 alerts per month. In February 2024, the model raised 14,540 alerts, which would ordinarily require around 121 individuals to process. However, by using the model this was dramatically reduced by the equivalent of approximately 45.6 full-time employees (FTE) (23.4 FTE for low-risk alerts and 22.2 FTE for high-risk alerts).
As business volumes increase further, the total alerts are also predicted to increase, and hence the number of FTEs saved is also forecast to increase. This underscores the significant impact the new model can continue to have on operational efficiency, potentially serving to help mitigate the increases in alerts generated by higher transaction volumes.
Discussion
This study had two aims: firstly, to outline the creation of a “bolt on”, ML hibernation model for managing money laundering transactions; Secondly, to examine the results of implementing the model in a real-world environment to assess the model’s efficiency. Both have been met and, in this section, we discuss the implications.
Technological and methodological innovation
We began this article by detailing how the risk scoring model was developed and implemented. In doing so, we have provided a blueprint for other financial institutions to consider similar approaches to AML. This is significant because previous studies of ML applications in the domain of AML (Alshantti and Rasheed, 2021; Desrousseaux et al., 2021; Gao, 2009; Jullum et al., 2020; Ketenci et al., 2021; Larik and Haider, 2011; Liu et al., 2008; Lv et al., 2008; Muhammed Shokry, 2020; Rui and Wunsch, 2005) have not explored this type of use before, particularly in the context of ML ensembled techniques (Oztas et al., 2022). It also addressed several cited limitations of previous research by increasing the depth of information in the data sets used and providing a formal evaluation (Oztas et al., 2022). Importantly, the blueprint and evaluation metrics used in this study ensure its reproducibility, addressing another previously identified issue (Oztas et al., 2022).
In our literature review we outlined how industry experts considered alternative methods to improve the efficiency of TM systems, such as segmentation and tuning, to be limited due to the need to regularly adjust thresholds, which resulted in high numbers of false positives (Oztas et al., 2024). This study demonstrated that ML models can complement this function effectively, relieving operational teams of significant work and creating additional capacity for other activities.
Another benefit of the study is confirmation of viability of the “bolt-on” approach. Industry experts have highlighted this as a way financial institutions can reap benefits of ML without abandoning traditional measures, such as the “rules-based” approach presently used (Oztas et al., 2024). Thus, financial institutions could integrate the models with assurance they will not only maintain but enhance the status quo. This should provide reassurance for compliance units and regulators that this approach is safe.
Efficiency improvements
In addition to technological and methodological contributions, this article demonstrated the models were highly efficient at categorizing suspicious activity. Their success in consistently identifying low-risk alerts correctly significantly reduced the volume of false positives that needed manual review. Prior to implementation, 100% of alerts had to be reviewed, by the end of the examined period, 19% of alerts were hibernated and 18% of alerts were auto-escalated resulting in a marked improvement in the effective alert rate at level 1, demonstrating that the “bolt-on” model performed significantly better than the TM systems alone.
Furthermore, this study identified that the model contributed to a significant reduction in time taken to investigate and submit SARs, which reduced by 61%. This greatly enhances the timeliness of investigations, which is crucial for money laundering prevention as research indicates that timely responses enable institutions and law enforcement agencies to make better use of the investigation timeframe after an offence is identified, enabling them to more rapidly address threats, freeze assets when necessary, and secure vital evidence before it is lost (Zolkaflil et al., 2019).
Limitations and further research
This study has several limitations that should be acknowledged and can be addressed through future research. Firstly, its scope is limited to a single major bank. This means that the model created was not validated across different financial institutions. As it was developed bespoke for the bank in question, this limits the generalizability of the findings, albeit we suggest the methodology can be universally applied. In this regard, application of the method in other AML areas, such as to assist in sanctions screening or monitoring of politically exposed persons (PEP) would broaden literature on the application of AI and ML in AML. Second, the reliance on historical transaction data and analysts’ decision-making may have introduced potential bias. This is a flaw inherent to all AI but one worthy of further examination to ensure the model does not unjustifiably or disproportionately impact any one demographic. Finally, we used a single ML methodology; therefore, incorporating other techniques, including deep learning, reinforcement learning and graph-based models, could enhance or improve predictive performance and capture more complex relationships in the data, and should therefore also be explored further in the context applied in this study.
Conclusion
Together, these findings echo previous studies that have outlined the potential for improved efficiency to reduce operational costs associated with AML compliance (Khan et al., 2018). The implications of these contributions are that they can help banking institutions effectively keep pace with the volume of transactions, which has a positive effect on reputation through enhancing their AML response (Nicknora, 2024).
In conclusion, improved efficiency is vital for banks as those institutes who are better at AML compliance also outperform competitor’s (McCarthy et al., 2019). This should serve as a commercial motivator for others to consider the blueprint we have presented.




