Training scenario (model type)
| Evaluation metric [test data] | All types (binary) | No non-reported alerts/cases (binary) | No normal transactions (binary) | All types (Multiresponse 3) | All types (Multiresponse 4) |
|---|---|---|---|---|---|
| AUC [all] | 0.907 [0.893, 0.921] | 0.852 [0.836, 0.869] | 0.872 [0.849,0.895] | 0.910 [0.896, 0.924] | 0.910 [0.895, 0.924] |
| AUC [only alerts] | 0.822 [0.795, 0.848] | 0.706 [0.675, 0.738] | 0.819 [0.791,0.847] | 0.826 [0.799, 0.854] | 0.825 [0.798, 0.853] |
| Brier [all] | 0.025 [0.022, 0.028] | 0.340 [0.330, 0.351] | 0.024 [0.021,0.028] | 0.025 [0.022, 0.027] | 0.025 [0.022, 0.027] |
| Brier [only alerts] | 0.047 [0.044, 0.051] | 0.655 [0.646, 0.665] | 0.047 [0.043,0.051] | 0.047 [0.044, 0.051] | 0.047 [0.044, 0.051] |
| PPP(TPR = 0.95) [all] | 0.315 | 0.373 | 0.452 | 0.305 | 0.330 |
| PPP(TPR = 0.8) [all] | 0.203 | 0.260 | 0.268 | 0.195 | 0.196 |
| Evaluation metric [test data] | All types (binary) | No non-reported alerts/cases (binary) | No normal transactions (binary) | All types (Multiresponse 3) | All types (Multiresponse 4) |
|---|---|---|---|---|---|
| AUC [all] | 0.907 [0.893, 0.921] | 0.852 [0.836, 0.869] | 0.872 [0.849,0.895] | 0.910 [0.896, 0.924] | 0.910 [0.895, 0.924] |
| AUC [only alerts] | 0.822 [0.795, 0.848] | 0.706 [0.675, 0.738] | 0.819 [0.791,0.847] | 0.826 [0.799, 0.854] | 0.825 [0.798, 0.853] |
| Brier [all] | 0.025 [0.022, 0.028] | 0.340 [0.330, 0.351] | 0.024 [0.021,0.028] | 0.025 [0.022, 0.027] | 0.025 [0.022, 0.027] |
| Brier [only alerts] | 0.047 [0.044, 0.051] | 0.655 [0.646, 0.665] | 0.047 [0.043,0.051] | 0.047 [0.044, 0.051] | 0.047 [0.044, 0.051] |
| PPP(TPR = 0.95) [all] | 0.315 | 0.373 | 0.452 | 0.305 | 0.330 |
| PPP(TPR = 0.8) [all] | 0.203 | 0.260 | 0.268 | 0.195 | 0.196 |
Notes:
Table I. Performance scores on the test set when training the model on all transactions (A-D), excluding normal transactions (A) and excluding non-reported alerts/cases (B + C) in the training. The two rightmost columns contain performance scores on the test set with a multiresponse with three or four outcomes. All models are trained on 13,782 transactions. The Brier and AUC scores are evaluated on test sets with either only alert data (ntest, only alerts = 2,557) or all data (alert data and normal transactions, ntest, all = 4,967). The Brier and AUC scores are given with 90 per cent CI in brackets
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.