Essay ·
Fraud, intrusion and a failing pump are the same problem.
Money laundering, account takeover, lateral movement and machine faults look like different industries. Underneath, they ask the same question of the data, which is why one engine can answer all of them.
A laundering ring moving money through pass-through accounts. An attacker with valid credentials hopping from host to host. A pump whose vibration is creeping away from its siblings. Different teams, different budgets, different vendors.
The question each of them asks of the data is the same: what does normal look like here, for this entity, among its peers, over time, and what does not fit?
Three kinds of anomaly
It helps to be precise about what “does not fit” means, because treating every outlier as a culprit is how alert queues fill with noise.
- Adversarial. Someone is working against you: laundering, fraud, intrusion. The source adapts once it is caught.
- Accidental. Something broke: a misconfigured job, a failing part, a retry loop duplicating payments.
- Artifact. The data is wrong while the behaviour is fine: a schema change, a timezone bug, a sensor recalibrated without anyone saying so.
The research literature adds a second axis. A single value can be abnormal on its own. A normal value can be abnormal in its context. Or every event can be individually normal while the pattern is not. Laundering and intrusion live mostly in the last two, which is why single-event rules miss them.
The same shapes, different names
Look past the vocabulary and the structures repeat.
- Money fanning out to many new counterparties looks like a user reaching many machines for the first time.
- A burst of small transfers just under a threshold looks like a burst of login attempts just under a lockout limit.
- A seller whose refund rate drifts from similar sellers looks like a machine whose readings drift from machines of the same type.
What changes between domains is the vocabulary, the typologies, and what a reviewer needs to see.
New knowledge, not a rewrite
That is how Lucir is built. Nothing in it is hard-coded to one dataset, and a new domain arrives as domain knowledge, not new code.
We proved it first where the evidence is hardest to argue with, on public anti-money-laundering benchmarks under a strict protocol. We then pointed the same engine at industrial predictive-maintenance data, and it built and certified detectors there with no dataset-specific code.
Security is next. Identities, hosts and sessions have the same shape as accounts, counterparties and transfers. If that is your world, we would like to hear from you.