The most dependable result is machine-learning triage placed after a bank's existing alert rules. At two banks it removed most false alerts while keeping about 90% of true cases.1,2 Network features added to standard tree models also work, and fast enough for real time.3,4 Three things still fail. Large language models make poor fraud classifiers.5,6 Transaction foundation models did not work for money laundering.7 Most models break when crime patterns or labels shift.8,9
Recommendation: a UK bank or payment firm should first build or buy mule-account scoring for incoming payments, with an explanation for each flag. Under UK scam rules, the receiving firm now pays half of each reimbursement.10 A live bank pilot of this kind of scoring beat the bank's rules.11
Caveat: most results are offline, from a single institution, and reported by the authors themselves.