Talk to us

Proof

Measured the way a sceptic would.

Three public benchmarks, no human in the build, scored on data the agent never saw.

The protocol

Three rules that keep the answer key locked.

  1. 01

    Never seen

    Each detector is scored on recent history the agent never saw.

  2. 02

    Never tuned on

    Scored when the detector registers, and never fed back into the build.

  3. 03

    Scored strictly

    Every alert must match one real transaction.

Benchmark

IBM AML, HI-Small

Transactions
5.08 million
Laundering
0.10%
Split
Temporal 60 / 20 / 20
Scored on
1,015,564 held-out transactions

Source: Altman et al., Realistic Synthetic Financial Transactions for Anti-Money Laundering Models, NeurIPS 2023

Minority-class F1, per transaction

0.557

precision 0.897 recall 0.404

About 5.5 hours, unattended · $4.63 in model calls

Published results on the same split, sorted by F1. Source: Altman et al..
Model F1
GIN 0.287
GIN + edge updates 0.477
Lucir, built by the agent 0.557
PNA 0.568
LightGBM + graph features 0.629
XGBoost + graph features 0.632

About a point below PNA, a purpose-built graph network. The best published methods are further ahead.

Precision 0.897: about nine in ten alerts are laundering.

Scoring: 0.584 at registration under our earlier, looser matching; 0.557 is the same detector re-scored strictly. Third unattended attempt; the first two registered nothing.

Benchmark

Elliptic Bitcoin

Transactions
203,769
Time steps
49
Split
Train 1 to 34, test 35 to 49
Features
166 per transaction

Source: Weber et al., Anti-Money Laundering in Bitcoin, KDD 2019 Workshop on Anomaly Detection in Finance

Illicit-class F1

0.625

precision 0.816 recall 0.506

8 minutes · $0.42 in model calls

Published results on the same split, sorted by F1. Source: Weber et al..
Model Precision Recall F1
Logistic regression, all features 0.404 0.593 0.481
Lucir, built by the agent 0.816 0.506 0.625
GCN 0.812 0.512 0.628
MLP, all features 0.694 0.617 0.653
Skip-GCN 0.812 0.623 0.705
EvolveGCN 0.850 0.624 0.720
Random forest, all features 0.956 0.670 0.788

Level with the paper's GCN, in eight minutes. The random forest is still well ahead.

A dark market closed at step 43; the paper reports no model catches what follows.

Benchmark

IEEE-CIS card fraud

Transactions
590,540 labelled
Fraud
3.5%
Split
Last 20% by time held out
Scored on
118,107 held-out transactions

Source: Vesta Corporation and IEEE-CIS, Fraud Detection competition, Kaggle 2019

ROC-AUC, every held-out transaction

0.926

PR-AUC 0.576 F1 0.550

61 minutes, unattended

Reference models we trained on the same split and scored the same way. Kaggle's leaderboard scored a hidden test set and is not comparable.
Model ROC-AUC PR-AUC F1
Plain gradient boosting, raw columns 0.920 0.582 0.553
1st-place recipe, past data only 0.924 0.583 n/a
Lucir, built by the agent 0.926 0.576 0.550
1st-place recipe, whole period 0.945 0.663 n/a

Level on ROC-AUC with the published winning recipe when that recipe may only use earlier data, as live monitoring must, and a little behind it on PR-AUC.

Most of the winners' lead came from aggregates computed over the whole period, future included, which Kaggle's rules allowed. What the leaderboard knew

A correction

We withdrew a number. Here is the one that holds up.

In August our internal score was 0.567. We then found three ways the holdout leaked into the build, closed them, and rebuilt: 0.584 on the same scoring.

Then we made the scoring strict, one flag to one transaction. That gives 0.557, the only figure we stand behind.

Read the full account

Next

A third benchmark, when it is honest.

Known typologies injected into real bank history. Published once injected cases leave no trace.

Reproduce it

Bring your own sceptic.

Same protocol, your dataset, your team watching.

Talk to us