Talk to us
← Notes

Essay ·

Writing detection rules is getting easy. Trusting them is the product.

Every monitoring vendor will soon have an agent that writes rules. The hard part, and the part regulated teams pay for, is proving those rules work and keeping them from drifting.

Ask a capable language model to write a transaction-monitoring rule and you will get one. Ask it twice and you will get two different rules. Both will look reasonable. Neither comes with evidence.

That is where the market is heading. Agentic rule-writing is quickly becoming table stakes, and every vendor in financial crime, fraud and security will offer some version of it. Generation is not the bottleneck any more. Trust is.

What a compliance team actually needs

A detector in a regulated setting decides who gets investigated. The questions that follow it are not about how it was written. They are about whether it works and whether it is still the thing that was approved:

  • Does it catch what it claims to, on data it was not tuned on?
  • How many false alarms will it raise, and can the team absorb them?
  • Is the version running today the version that was approved?
  • When an examiner points at one alert, can we show exactly why it fired?

Model risk guidance in banking, such as the Federal Reserve’s SR 11-7, has long asked for validation that is independent of development. An agent that both builds a detector and grades it cannot offer that, however good the agent is. Separating the two inside the product does not replace a bank’s own validation, but it gives that validation something solid to check.

Separate the builder from the judge

So we split them. The agent builds. Certification checks the result, and a final score comes from history the agent never sees. A person approves, and the approved logic is frozen.

  • Evidence before deployment. Accuracy and flag-rate checks, a review against the original intent, and a score on history the agent never saw.
  • Frozen once approved. The logic stays as approved, and retrained models are versioned.
  • The same answer twice. A certification can be re-run and checked.

Why this is the moat

Generation will keep getting cheaper and better, for everyone. Rigor does not come free with a better model. It has to be designed into the system that surrounds the model: what the agent can see, what it is judged on, and what happens after approval.

That system is what we are building. The agent is how it gets fast. The evidence is why anyone should use it.