Engineering note ·
We withdrew our best number. Here is the one that holds up.
We found three ways our test set was leaking into the build, withdrew the result, fixed the protocol, then made the scoring stricter too. The number we quote now is lower, and it is the one worth quoting.
In August our internal benchmark said Lucir’s agent had built a detector scoring 0.567 on IBM’s HI-Small anti-money-laundering benchmark, measured as minority-class F1. That put it next to PNA, a graph neural network designed for exactly this task. We were pleased with it.
Then we reviewed our own certification code the way an examiner would, and found three ways the held-out test data was reaching the build.
Three leaks
The agent could read it. After every test run, the report the agent received included the holdout score, and the build used it to choose between candidates. Choosing between candidates by their test score quietly turns the test set into a validation set, and the number that comes out is optimistic.
It was a gate. The holdout score could pass a detector against the published floor, or fail one that fell below it. Either way, test labels were steering what survived.
Models trained on it. Models were fit on labeled rows that included the holdout, and tuned on a slice inside it.
None of this was deliberate, and each one is easy to miss.
What we changed
We withdrew the number and rebuilt the protocol around one idea: the agent that builds a detector must never be able to see, or be steered by, the data that judges it. The test is scored once, when the detector registers, and never used to tune anything.
Then we ran the benchmark again, from scratch, with no one touching the build.
A surprise, and a second fix
The clean build scored 0.584 under the same scoring as before. Higher, not lower. The leaks had been real, but they were not what made the old number look good.
That sent us to the scoring itself. It matched each flag to transactions by sender and time, which is generous when one sender makes several transfers in the same moment. We made it strict: every flag is matched to exactly one source transaction, and an ambiguous match counts against us. Re-scoring the same registered detector that way, with no rebuild, gives 0.557, at a precision of 0.897. Roughly nine in ten alerts are laundering.
PNA sits at 0.568. The best result in the benchmark’s own paper, gradient boosting with graph features, is 0.632, and later work reports higher still. The full table is on our proof page.
The question to ask any vendor
A detection score is only as good as the protocol and the scoring behind it. So ask, of us and of anyone else:
Who chose the model, what could it see while choosing, when was the test scored, and how is a flag matched to the thing it flags?
We would rather publish a lower number that survives those questions.