Engineering note ·
What the leaderboard knew
We rebuilt the winning recipe from a famous fraud competition on our own time split. Most of its lead came from averages that already include each card's later transactions, which a live system never has.
In 2019, 6,351 teams competed on Vesta’s card-not-present fraud data, the IEEE-CIS Fraud Detection competition on Kaggle. The winning team finished at 0.9459 ROC-AUC on the private leaderboard, and its write-up became a standard reference for fraud feature engineering.
We added the dataset to our benchmarks. Kaggle never released the test labels, so we hold out the last fifth of the labelled data by time, as we do for every benchmark, and score every held-out transaction.
A yardstick we built ourselves
No published result uses our split, so we built the comparison ourselves. We reimplemented the published first-place single model: its reconstructed card identity, its per-card averages and counts, and its model settings. We trained it on exactly the data Lucir sees and scored it with the same code.
On our holdout it scores 0.945. Close to its reputation.
Where the lead comes from
The competition’s rules let entrants compute features over the training and test data together, and the winning recipe does. So when it scores a transaction, “this card’s average amount” and “how many times this card appears” already include that card’s later transactions, often the rest of a fraud spree.
We changed one thing: every average and count may only use events that happened earlier, as it would have to in production. The same recipe then scores 0.924. A plain gradient-boosted model on the raw columns scores 0.920.
So the part of the recipe a monitoring system can actually run is worth about half a point. The other two points came from knowing the future.
Where Lucir stands
Lucir builds its detectors unattended and watches data as it arrives, so everything it learns from reads only the past, by design. Given one sentence, “Flag fraudulent transactions”, its latest build on this benchmark scores 0.926. That is level with the live-safe version of the winning recipe (0.924) and a little behind it on precision-recall (0.576 against 0.583). The remaining gap to 0.945 is the part that needed the future.
This is not a criticism of the winners. They played a competition by its rules, and played it brilliantly. But a leaderboard number is only as useful as the conditions behind it. When a vendor quotes one, ask a simple question: could each of these predictions have been made at the moment the transaction happened?