Between Papers

A published investigation. Shown in full, as it was produced. Open any citation to see the passage it rests on; the method explains what was checked. Start an investigation on your own question.

Investigation report, 29 Sept 2026

AI against payment fraud and financial crime

What have recent advances in AI for payment fraud and financial-crime detection made reliable, what still fails, and what products could a startup or bank build now?

The answer

The most dependable result is machine-learning triage placed after a bank's existing alert rules. At two banks it removed most false alerts while keeping about 90% of true cases.1,2 Network features added to standard tree models also work, and fast enough for real time.3,4 Three things still fail. Large language models make poor fraud classifiers.5,6 Transaction foundation models did not work for money laundering.7 Most models break when crime patterns or labels shift.8,9

Recommendation: a UK bank or payment firm should first build or buy mule-account scoring for incoming payments, with an explanation for each flag. Under UK scam rules, the receiving firm now pays half of each reimbursement.10 A live bank pilot of this kind of scoring beat the bank's rules.11

Caveat: most results are offline, from a single institution, and reported by the authors themselves.

How we got here

  1. Started from 25 seed papers

    • Anti-Money Laundering Alert Optimization Using Machine Learning with Graphs 2021
    • Turning the Tables: Biased, Imbalanced, Dynamic Tabular Datasets for ML Evaluation 2022
    • Realistic Synthetic Financial Transactions for Anti-Money Laundering Models 2023
    • Finding Money Launderers Using Heterogeneous Graph Neural Networks 2023
    and 21 more
    • Searching for Smurfs: Testing if Money Launderers Know Alert Thresholds 2023
    • Topology-Agnostic Detection of Temporal Money Laundering Flows in Billion-Scale Transactions 2023
    • Graph Feature Preprocessor: Real-time Subgraph-based Feature Extraction for Financial Crime Detection 2024
    • The Shape of Money Laundering: Subgraph Representation Learning on the Blockchain with the Elliptic2 Dataset 2024
    • Graph Neural Networks for Financial Fraud Detection: A Review 2024
    • Year-over-Year Developments in Financial Fraud Detection via Deep Learning: A Systematic Literature Review 2025
    • Towards Collaborative Anti-Money Laundering Among Financial Institutions 2025
    • Advances in Continual Graph Learning for Anti-Money Laundering Systems: A Comprehensive Review 2025
    • AMLGENTEX: MOBILIZING DATA-DRIVEN RESEARCH TO COMBAT MONEY LAUNDERING 2025
    • Exploring the In-Context Learning Capabilities of LLMs for Money Laundering Detection in Financial Graphs 2025
    • 'Send to which account?' Evaluation of an LLM-based Scambaiting System 2025
    • Spatio-Temporal Directed Graph Learning for Account Takeover Fraud Detection 2025
    • PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling 2025
    • Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions 2025
    • TREASURE: The Visa Payment Foundation Model for High-Volume Transaction Understanding 2025
    • Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detection 2025
    • REINFORCEMENT LEARNING OF LARGE LANGUAGE MODELS FOR INTERPRETABLE CREDIT CARD FRAUD DETECTION 2026
    • PRAGMA: Revolut Foundation Model 2026
    • HELPING CUSTOMERS IN DISTRESS: AN LLMPOWERED AGENT THAT CONVERSES, PROBES, AND ROUTES 2026
    • Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification 2026
    • MINT: A Universal Zero-Shot Predictor for Transaction Data 2026
  2. Looked at 266 sources, read 5 in full

    1. Looked at266
    2. Screened266
    3. Shortlisted50
    4. Read in full5

    Works citing the seeds 47 · Related works 38 · Business and economics papers 24 · Adjacent-topic search 50 · Industry whitepapers (web search) 107

    Every source was screened from its title and abstract before any was read; 30 shortlisted sources could not be fetched (no open copy or unreadable). 25 verified claims came from the sources read.

    What we read

    Related papers 1

    Industry whitepapers 4

  3. Verified 172 claims

    240 checks run against the source passages; 46 claims caught and corrected on the way.

    Distilled into 6 findings, 4 opportunities, 3 combinations.

What we found

Capability

ML triage on top of AML rules cuts false alerts sharply, at two banks

A model that re-scores rule alerts removed 80% of false positives while keeping over 90% of true positives at one bank. An independent study at a second bank cut false positives by 0.76 (as a share) at 0.90 recall. Both are offline tests on alert sets dominated by false positives.1,2+3

Capability

Network features matter more than the choice of graph model

Adding subgraph features to boosted trees beat graph neural networks by up to 8% on synthetic AML data, and small-batch CPU latency averaged 30 ms. On real Capital One data a graph model beat production XGBoost; the authors credit the graph formulation, not model depth.3,4+3

Contradiction

Transaction foundation models lift fraud recall but fail at money laundering

WeChat Pay, Revolut and Visa report large relative fraud gains over their production models, one in a live A/B test. Revolut's model scored 47.1% below its AML baseline, blamed on missing cross-account network features. All figures are relative and unreproducible.7,18+2

Limit

LLMs are poor fraud detectors but useful explainers and triage agents

Untuned frontier LLMs scored low F1 on real e-commerce fraud, and even retrieval-aided LLMs trail a tabular model. LLM-written mule-case narratives took seconds per alert on one GPU and a bank triage agent beat its legacy flow, but both need human review.5,6+5

Limit

Models break under drift, rarer crime and noisy labels

A model trained on one synthetic laundering mix scored F1 near 0 on another until retrained. Temporal shift degraded fraud models, labels are sparse and noisy at real banks, and federated training across banks helped only slightly.8,9+4

Cost

Compliance cost and new UK liability rules create paying buyers

LexisNexis/Forrester put global financial-crime compliance cost at $206B, with salaries the main driver in most regions. UK rules now make receiving banks fund 50% of scam reimbursements up to £85,000, and the regulator says no bank detects all such fraud.10,30+3

What you could build

Each is rated on evidence, technical readiness and commercial access separately; there is no overall score. Hover a rating, or open “Why these ratings”, for the reasoning.

Alert-triage layer with auto-drafted case narratives for AML teams

BankingPaymentsRegTechProduct
EvidenceHighReadinessHighAccessMedium

A scoring layer that sits after a bank's existing monitoring rules, ranks or auto-closes low-risk alerts, and drafts an explanation of each score for the analyst. Rules stay in place, which keeps the audit trail regulators expect.1,2+6

Who buys
Heads of financial-crime compliance or transaction monitoring at banks and fintechs, bought as SaaS or an on-premise licence priced per alert volume.
Why now
Two separate banks show large false-positive cuts at high recall, and a live bank pilot lifted alert yield with LLM-written narratives on a single GPU. Salaries are the main compliance cost driver in most regions.
Why these ratings
Evidence: High
Replicated at two banks with similar results, plus a short live pilot at another bank; all authors report their own results.
Readiness: High
Uses boosted trees and explanation methods banks already run; narration fits on one on-premise GPU.
Access: Medium
Crowded vendor market and long bank procurement; model-risk approval needed before auto-closing alerts.
Assumptions and main risk
  • Offline false-positive cuts hold on live alert streams at natural class balance
  • Regulators accept model-closed alerts when rules remain the first line
  • Banks have enough historical investigator labels to train per client

Main risk. Narratives that miss facts, and labels built from past investigator decisions, can hide missed laundering and draw regulatory criticism.

Receiving-side mule account detection for UK payment firms

BankingPaymentsE-moneyProduct
EvidenceMediumReadinessHighAccessHigh

A mule-risk score for incoming Faster Payments accounts that combines account behaviour with network features and explains each flag, so the receiving firm can hold or freeze funds before they move on.4,10+5

Who buys
Fraud and financial-crime leads at UK banks, EMIs and payment institutions that now pay half of every APP scam reimbursement as receiving firms.
Why now
Since October 2024 receiving firms fund half of reimbursements up to £85,000, a direct cost for hosting mule accounts. A bank's mule model with explanations already outperformed its rules in a live pilot.
Why these ratings
Evidence: Medium
One live mule-detection pilot at one bank; liability rules are primary regulator text, but loss offsets are unmeasured.
Readiness: High
Boosted-tree classifiers and fast graph features already run at real-time latency.
Access: High
Every UK receiving firm now carries a new, countable liability; smaller EMIs lack in-house data science.
Assumptions and main risk
  • Mule models transfer across institutions after retraining
  • Reimbursement costs are large enough per firm to justify spend
  • Holding flagged incoming funds is operationally and legally acceptable

Main risk. Fraud may pivot to international payments outside the scheme, shrinking the liability that motivates buyers.

Real-time graph-feature engine for existing fraud models

PaymentsBankingCrypto complianceComponent
EvidenceMediumReadinessMediumAccessMedium

A streaming service that computes fan-in, fan-out, cycle and other subgraph features per transaction and feeds them into the boosted-tree models banks already run, without replacing those models.3,4+5

Who buys
Fraud-model and ML platform teams at PSPs, card issuers and banks, or vendors embedding it, paid per transaction volume.
Why now
Graph features gave large accuracy gains to tree models and beat graph neural networks on synthetic AML data at real-time CPU latency, while single-GPU graph networks do not reach bank scale.
Why these ratings
Evidence: Medium
Strong gains, but mostly on synthetic data; real-bank graph evidence comes from a different method.
Readiness: Medium
CPU latency is measured; real-time use caps batch size and costs some accuracy.
Access: Medium
Fits current tree-model stacks, but large banks may build in-house.
Assumptions and main risk
  • Synthetic-data gains carry over to real transaction graphs
  • Buyers can supply streaming counterparty data within latency budgets

Main risk. Gains measured on synthetic laundering patterns may shrink on real data where launderers blend in.

Stress-testing service for fraud and AML model validation

BankingRegTechModel riskService
EvidenceMediumReadinessHighAccessLow

A validation kit that runs a bank's model against synthetic scenarios: rarer crime, new laundering mixes, time drift, scarce or noisy labels, and group fairness. It outputs a model-risk report.8,26+6

Who buys
Model risk management and internal audit teams at banks, and vendors needing evidence for client procurement.
Why now
Open generators now simulate label noise, drift and partial visibility, and they show models collapsing when the laundering mix changes. Reviews note no shared benchmarks with operational metrics.
Why these ratings
Evidence: Medium
Failure modes are well shown, but only on synthetic data; the gap to real data is not measured.
Readiness: High
Generators and benchmark suites are public and already used in research.
Access: Low
Budget sits in model-risk teams; willingness to pay for synthetic tests is unproven.
Assumptions and main risk
  • Regulators or auditors value synthetic stress tests as evidence
  • Synthetic failure modes predict real-world failures

Main risk. Without proof that synthetic results predict real performance, buyers may treat the report as a checkbox.

Ideas from combining sources

Two concepts from different sources, tied into an idea neither states. These are hypotheses to test, not findings.

Scambaiting bots as a fresh label source for mule-account models

BankingPaymentsFraud intelligence
Concept A

An LLM scambaiting system engaged real scammers and extracted financial details such as mule accounts in 31.74% of conversations where scammers replied.47,48,49

From 'Send to which account?' Evaluation of an LLM-based Scambaiting System
Concept B

A bank's mule classifier works in production, but its labels come from police filings and analyst confirmations that arrive late and are biased.11,29

From Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification
Together

Feed mule accounts harvested by scambaiting agents, shared through industry channels, into receiving banks' mule models as early positive labels and watchlist hits. This targets the label delay that weakens today's mule models, and catches accounts before any victim reports them.

The scambaiting work stops at intelligence sharing; the mule-detection work relies on slow police and analyst labels. Neither links the two.

Where it breaks
Harvested accounts may be few, short-lived or already burned; only 48.7% of engagements got any reply, and label provenance may not satisfy bank legal teams.
How to test it
Match a month of harvested account numbers against a bank's accounts; measure how many were later confirmed as mules and how many days earlier they were flagged than police-derived labels.

Foundation-model embeddings fused with streaming graph features for AML

BankingFintechRegTech
Concept A

A bank transaction foundation model failed on AML, blamed on missing cross-account network features; its kind of embeddings can be cached and fused live.7,50

From PRAGMA: Revolut Foundation Model
Concept B

A graph-feature preprocessor computes subgraph patterns per transaction at real-time latency and lifts tree models sharply on laundering data.4,38

From Graph Feature Preprocessor: Real-time Subgraph-based Feature Extraction for Financial Crime Detection
Together

Build the AML model as a tree model that takes both per-customer foundation-model embeddings (behaviour over time) and streaming subgraph features (who pays whom). Each supplies what the other lacks: behavioural history versus network structure.

Foundation-model papers leave graph modelling as future work, and graph-feature work uses no learned behavioural embeddings.

Where it breaks
Evidence for each part is from different data (proprietary bank vs synthetic); the fusion's gain is untested and labels for AML stay scarce.
How to test it
On one bank's AML labels, compare embeddings alone, graph features alone and both combined in the same tree model, at a fixed alert budget.

Privacy-preserving sketch exchange between sending and receiving banks

BankingPaymentsRegTech
Concept A

Banks can find cross-institution laundering groups by sharing only hashed counterparty summaries, not raw transactions; tested on real Alipay data.51,52,53

From Towards Collaborative Anti-Money Laundering Among Financial Institutions
Concept B

UK rules make receiving firms pay half of each APP scam reimbursement, explicitly to give them an incentive to detect the mule accounts they host.10,33

From PS25/5 APP scams reimbursement requirement consolidated policy statement
Together

Offer a shared service where sending and receiving firms exchange hashed counterparty sketches at payment time, so both sides can spot mule clusters that span firms without breaching data-protection rules. Shared liability gives both sides a reason to join.

The cross-bank method was pitched for AML research, not tied to a liability rule that pays both parties to participate.

Where it breaks
The method was tested on two institutions and one laundering pattern; payment-time latency and formal privacy leakage are unmeasured.
How to test it
Pilot with two UK firms on past APP scam cases: check whether sketch matching flags confirmed mule receivers earlier than each firm alone.

Evidence appendix

The detailed analysis behind the cards above, section by section.

Summary of the investigation9 claims cited

Question. The prompt supplied this question. It asks what AI for payment fraud and financial-crime detection now does reliably, what still fails, and what a startup or bank could build now.

Evidence base. We read all the seed papers. They cover graph methods for anti-money-laundering (AML), synthetic data and benchmarks, transaction foundation models, and LLM agents. We screened a further set of candidate sources and read five of them:

  • the UK Finance annual fraud report, via a law-firm summary;54
  • the UK payments regulator's (PSR) scam-reimbursement policy;10
  • the Bank of England/FCA survey of AI use;55
  • the LexisNexis/Forrester compliance-cost study;30
  • an independent bank study on suppressing false AML alerts, which replicates a seed's result.2

Strength. The evidence on alert triage is strong, because two banks report similar results. Other evidence is moderate:

  • Graph features were tested mostly on synthetic data.56
  • Foundation-model gains are relative figures against undisclosed internal baselines.57
  • Several real-data studies have no labelled ground truth, or check only small samples.58,59

Not reached.

  • The FATF new-technologies report exceeded the page cap, and its mirror timed out. The global regulator's stance on explainability is therefore covered only through the UK survey.
  • The OpenAlex search for grey-literature reports failed, and several seeds were not found in OpenAlex.
  • No source measured money lost to fraud that was prevented, or analyst hours saved.
  • We did not examine card-network or vendor production benchmarks, or evidence on LLM-drafted suspicious-activity reports beyond one pilot.
What the sources agree on20 claims cited
  • Rule alerts are mostly noise, and ML triage fixes much of it. More than 97% of alerts were false positives at one bank.13 A second bank's alerts were likewise dominated by false positives.14 Triage models cut false positives substantially at high recall at both banks.1,2
  • Tree models remain the workhorse. Boosted trees with graph features matched or beat graph neural networks on synthetic AML data.3,15 Production baselines at Capital One and OCBC were also tree models.16,60
  • Imbalance and scarce labels are the norm. Most reviewed deep-learning fraud papers report class imbalance.61 Real bank labels are sparse and noisy.28,29,62
  • Explainability is required, not optional. Reviews and model authors name it as an unmet need for regulated use.63,64,65
  • Humans stay in the loop. Only 2% of UK financial-firm AI use cases are fully autonomous.66 LLM narration and triage agents both needed human review.22,25
  • Adoption is mainstream. A third of UK firms already use AI for fraud detection, and firms rate AML and fraud among the top benefits.55,67
Where the evidence conflicts14 claims cited
  • Graph neural networks versus trees. On synthetic AML data, trees with graph features beat graph neural networks.3,15 On real account-takeover data, a graph model beat production XGBoost.16 A heterogeneous graph network also beat a plain network on real DNB data.68 The two sides agree on one point: the graph information matters more than the learner. The Capital One authors say deeper per-row models only matched XGBoost.17
  • Foundation models: fraud versus AML. Revolut's model improved fraud recall but fell 47.1% below the AML baseline.7,19 So a single "universal" transaction model is not yet universal.
  • LLMs on tabular fraud. With retrieval, LLMs beat Random Forest and XGBoost on one dataset.69 They still trailed a modern tabular model on others.6 After RL tuning, the larger model's explanations became less faithful.21
  • Do launderers game thresholds? A bunching test on real Danish bank data found no excess just below the alert thresholds.70 That undercuts the common assumption that criminals know the rules, though it is one bank.
  • Federated versus pooled data. Federated training across 12 simulated banks helped only slightly.9 A hashing-based cross-bank method did find real laundering groups, but without a baseline.52,59
What is now possible, and what limits it24 claims cited
  • Reliable now: triage after rules. This keeps rule explainability and removes most false alerts.2,12 Limit: it needs historical investigator labels, and it runs alongside rules, not instead of them.34
  • Reliable now: real-time graph features on CPUs. Latency is 30 ms for small batches, with throughput above GPU graph networks.4 Limit: smaller batches cost accuracy.39 Single-GPU graph networks reach only about ten million transactions.40
  • Emerging: flow detection at billion scale. It was run on over 1 billion real transactions.71 Limit: there is no ground truth, so no precision or recall is reported.58
  • Emerging: transaction foundation models at large platforms. They were pretrained on billions of events,72,73,74 and one was A/B-tested in production.18 Limit: results are relative only, they did not work for AML, and they are not interpretable.7,65
  • Emerging: LLM narration and conversational agents. Case narratives ran on one on-premise GPU.23 A bank's triage agent kept handoff precision and recall above 90%.75 A scambaiting agent extracted mule details.47 Limit: narratives are incomplete,22 and about half of scam engagements got no reply.48
  • Not reliable: LLMs as detectors. Their F1 is low untuned and modest after RL tuning.5,76 On synthetic AML graphs, accuracy was 63.7%.77
  • Not reliable: robustness to shift. Performance collapsed when the laundering mix changed,8 and fell under temporal shift.27
Who feels the pain today14 claims cited
  • AML compliance teams. LexisNexis/Forrester estimate global financial-crime compliance cost at $206B. Salaries are the main driver of cost increases in most regions.30,31 Alert triage and narrative drafting address this labour directly.1,11
  • UK receiving banks and payment firms. Since 7 October 2024, receiving firms pay 50% of APP scam reimbursements.10,35 The maximum reimbursement is £85,000.32 UK Finance reports APP losses of just over £450 million in 2024.78
  • Card and account fraud teams. UK Finance reports unauthorised fraud losses of £722 million.54 The relevant tools are account-takeover and card-fraud models that use graph or sequence signals.16,18
  • Customer-service teams handling scam victims. An LLM agent classified fraud and dispute reports better than a legacy flow.24
  • Model risk and fairness reviewers. Fraud models blind to fairness flagged legitimate older applicants far more often.44 Reviews call for benchmarks with operational metrics.45

Sources

Every citation marker opens the passage it rests on in one of these. Sources the agent added show why, and web pages show where and when they were read.

Seed papers 25

  • 00Anti-Money Laundering Alert Optimization Using Machine Learning with Graphs
    2021 · 8 pp106 spans over 8/8 pages, 36,604 chars, 106 with section paths, labels={'text': 80, 'section_header': 15, 'caption': 6, 'list_item': 3, 'footnote': 1, 'table': 1}
    clean
  • 01Turning the Tables: Biased, Imbalanced, Dynamic Tabular Datasets for ML Evaluation
    2022 · 19 pp206 spans over 19/19 pages, 66,909 chars, 206 with section paths, labels={'list_item': 88, 'text': 74, 'section_header': 28, 'caption': 9, 'table': 4, 'footnote': 3}
    clean
  • 02Realistic Synthetic Financial Transactions for Anti-Money Laundering Models
    2023 · 24 pp272 spans over 24/24 pages, 78,876 chars, 272 with section paths, labels={'text': 114, 'list_item': 87, 'section_header': 28, 'caption': 24, 'table': 15, 'footnote': 4}
    clean
  • 03Finding Money Launderers Using Heterogeneous Graph Neural Networks
    2023 · 20 pp215 spans over 20/20 pages, 67,715 chars, 215 with section paths, labels={'text': 122, 'list_item': 47, 'section_header': 27, 'caption': 10, 'footnote': 5, 'table': 4}
    clean
  • 04Searching for Smurfs: Testing if Money Launderers Know Alert Thresholds
    2023 · 14 pp149 spans over 14/14 pages, 39,972 chars, 148 with section paths, labels={'text': 55, 'list_item': 52, 'section_header': 22, 'caption': 10, 'footnote': 8, 'table': 2}
    clean
  • 05Topology-Agnostic Detection of Temporal Money Laundering Flows in Billion-Scale Transactions
    2023 · 19 pp · 10.1007/978-3-031-74643-7_291 items dropped for having no page provenance -- a span without a page a reader can turn to is not evidence; 157 spans over 19/19 pages, 44,906 chars, 157 with section paths, labels={'text': 79, 'list_item': 43, 'section_header': 17, 'caption': 14, 'footnote': 2, 'table': 2}
    clean
  • 06Graph Feature Preprocessor: Real-time Subgraph-based Feature Extraction for Financial Crime Detection
    2024 · 11 pp · 10.1145/3677052.3698674185 spans over 11/11 pages, 66,477 chars, 181 with section paths, labels={'list_item': 91, 'text': 56, 'caption': 16, 'section_header': 12, 'table': 5, 'footnote': 3, 'title': 1, 'other': 1}
    clean
  • 07The Shape of Money Laundering: Subgraph Representation Learning on the Blockchain with the Elliptic2 Dataset
    2024 · 7 pp134 spans over 7/7 pages, 41,390 chars, 132 with section paths, labels={'text': 61, 'list_item': 35, 'section_header': 23, 'caption': 6, 'table': 4, 'footnote': 4, 'title': 1}
    clean
  • 08Graph Neural Networks for Financial Fraud Detection: A Review
    2024 · 17 pp · 10.1007/s11704-024-40474-y2 items dropped for having no page provenance -- a span without a page a reader can turn to is not evidence; 313 spans over 17/17 pages, 77,952 chars, 313 with section paths, labels={'list_item': 170, 'text': 102, 'section_header': 32, 'caption': 5, 'table': 3, 'footnote': 1}
    clean
  • 09Year-over-Year Developments in Financial Fraud Detection via Deep Learning: A Systematic Literature Review
    2025 · 21 pp2 items dropped for having no page provenance -- a span without a page a reader can turn to is not evidence; 221 spans over 21/21 pages, 52,789 chars, 218 with section paths, labels={'text': 94, 'list_item': 82, 'section_header': 33, 'caption': 10, 'title': 1, 'table': 1}
    clean
  • 10Towards Collaborative Anti-Money Laundering Among Financial Institutions
    2025 · 12 pp · 10.1145/3696410.3714576198 spans over 12/12 pages, 61,627 chars, 196 with section paths, labels={'text': 103, 'list_item': 39, 'section_header': 30, 'caption': 11, 'table': 6, 'footnote': 6, 'other': 2, 'title': 1}
    clean
  • 11Advances in Continual Graph Learning for Anti-Money Laundering Systems: A Comprehensive Review
    2025 · 27 pp · 10.1002/wics.70040307 spans over 27/27 pages, 91,657 chars, 307 with section paths, labels={'text': 147, 'list_item': 93, 'section_header': 33, 'caption': 25, 'table': 5, 'footnote': 4}
    clean
  • 12AMLGENTEX: MOBILIZING DATA-DRIVEN RESEARCH TO COMBAT MONEY LAUNDERING
    2025 · 29 pp243 spans over 29/29 pages, 83,179 chars, 239 with section paths, labels={'text': 108, 'list_item': 55, 'caption': 37, 'section_header': 31, 'table': 9, 'footnote': 2, 'title': 1}
    clean
  • 13Exploring the In-Context Learning Capabilities of LLMs for Money Laundering Detection in Financial Graphs
    2025 · 4 pp · 10.1109/ICDMW69685.2025.0002662 spans over 4/4 pages, 19,165 chars, 57 with section paths, labels={'text': 33, 'section_header': 12, 'list_item': 9, 'caption': 5, 'table': 2, 'title': 1}
    clean
  • 14'Send to which account?' Evaluation of an LLM-based Scambaiting System
    2025 · 13 pp182 spans over 13/13 pages, 61,838 chars, 170 with section paths, labels={'text': 120, 'list_item': 26, 'section_header': 20, 'caption': 10, 'table': 5, 'title': 1}
    clean
  • 15Spatio-Temporal Directed Graph Learning for Account Takeover Fraud Detection
    2025 · 7 pp93 spans over 7/7 pages, 21,478 chars, 92 with section paths, labels={'text': 47, 'list_item': 23, 'section_header': 17, 'caption': 3, 'title': 1, 'footnote': 1, 'table': 1}
    clean
  • 16PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling
    2025 · 22 pp341 spans over 22/22 pages, 70,815 chars, 336 with section paths, labels={'text': 146, 'list_item': 121, 'section_header': 50, 'caption': 12, 'table': 9, 'footnote': 2, 'title': 1}
    clean
  • 17Open Banking Foundational Model: Learning Language Representations from Few Financial Transactions
    2025 · 5 pp · 10.1145/nnnnnnn.nnnnnnn79 spans over 5/5 pages, 27,554 chars, 77 with section paths, labels={'text': 39, 'section_header': 14, 'list_item': 11, 'caption': 6, 'footnote': 4, 'table': 4, 'title': 1}
    clean
  • 18TREASURE: The Visa Payment Foundation Model for High-Volume Transaction Understanding
    2025 · 10 pp157 spans over 10/10 pages, 53,933 chars, 155 with section paths, labels={'text': 71, 'list_item': 41, 'section_header': 22, 'caption': 14, 'table': 5, 'footnote': 2, 'title': 1, 'other': 1}
    clean
  • 19Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detection
    2025 · 16 pp203 spans over 16/16 pages, 68,009 chars, 198 with section paths, labels={'list_item': 77, 'text': 71, 'section_header': 24, 'caption': 14, 'table': 10, 'footnote': 6, 'title': 1}
    clean
  • 20REINFORCEMENT LEARNING OF LARGE LANGUAGE MODELS FOR INTERPRETABLE CREDIT CARD FRAUD DETECTION
    2026 · 18 pp1 items dropped for having no page provenance -- a span without a page a reader can turn to is not evidence; 161 spans over 18/18 pages, 60,950 chars, 158 with section paths, labels={'text': 67, 'list_item': 59, 'section_header': 19, 'caption': 11, 'footnote': 2, 'table': 2, 'title': 1}
    clean
  • 21PRAGMA: Revolut Foundation Model
    2026 · 20 pp213 spans over 20/20 pages, 76,484 chars, 208 with section paths, labels={'text': 93, 'list_item': 62, 'section_header': 40, 'caption': 10, 'table': 7, 'title': 1}
    clean
  • 22HELPING CUSTOMERS IN DISTRESS: AN LLMPOWERED AGENT THAT CONVERSES, PROBES, AND ROUTES
    2026 · 5 pp1 items dropped for having no page provenance -- a span without a page a reader can turn to is not evidence; 50 spans over 5/5 pages, 16,935 chars, 47 with section paths, labels={'text': 23, 'list_item': 12, 'section_header': 11, 'caption': 2, 'title': 1, 'table': 1}
    clean
  • 23Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification
    2026 · 10 pp164 spans over 10/10 pages, 52,296 chars, 162 with section paths, labels={'text': 78, 'list_item': 35, 'section_header': 28, 'caption': 9, 'table': 6, 'other': 4, 'footnote': 3, 'title': 1}
    clean
  • 24MINT: A Universal Zero-Shot Predictor for Transaction Data
    2026 · 8 pp147 spans over 8/8 pages, 47,715 chars, 144 with section paths, labels={'text': 85, 'list_item': 34, 'section_header': 16, 'caption': 5, 'table': 3, 'other': 2, 'title': 1, 'footnote': 1}
    clean

Related papers 1

  • 29Automatic suppression of false positive alerts in anti-money laundering systems using machine learning
    link.springer.com (opens in a new tab) · retrieved 29 Sept 2026 · found via OpenAlexCited-by P000: an independent ML false-positive suppression study on bank AML alerts, to test whether P000's alert-reduction result replicates elsewhere.
    web

Industry whitepapers 4

  • 25service.betterregulation.com/sites/default/files/2025-06/UK%20Finance%20Annual%20Fraud%20Report%202025.pdf
    service.betterregulation.com (opens in a new tab) · retrieved 29 Sept 2026 · found via ExaPrimary UK industry data on 2024 fraud losses, APP scam volumes and prevention rates: sizes the buyer pain for payment-fraud products.
    web
  • 26PS25/5 APP scams reimbursement requirement consolidated policy statement
    psr.org.uk (opens in a new tab) · retrieved 29 Sept 2026 · found via ExaUK regulator's APP scam reimbursement rules (liability split between sending and receiving PSPs, cap): the regulatory driver that makes scam and mule detection a paid-for problem.
    web
  • 27Artificial intelligence in UK financial services - 2024
    bankofengland.co.uk (opens in a new tab) · retrieved 29 Sept 2026 · found via Exa and TavilyBank of England/FCA 2024 survey of AI adoption, use cases (including fraud and AML) and governance constraints in UK financial firms: evidence on buyer adoption and regulator stance.
    web
  • 28cartographai.com/wp-content/uploads/2025/07/LNRS-TLP-True-Cost-Of-Financial-Crime-Compliance_2023.pdf
    cartographai.com (opens in a new tab) · retrieved 29 Sept 2026 · found via ExaLexisNexis/Forrester study of financial-crime compliance spend and its drivers (labour vs technology): sizes the AML compliance-cost market that alert-triage tools address.
    web

Run this on your own papers

Start with a few seed papers and a question. The free preview gives you the answer and a couple of findings; the full report is the same run, continued.

$149for the full report
Get a free preview