
Anomaly Detection - The Swiss Army Knife of Fraud Prevention
Anomaly detection spots what differs from the expected pattern, but unusual does not always mean fraud. Explore how baselines, context and combinations of signals help separate suspicious activity from normal behavior.
A payment clears authentication. Confirmation of Payee returns a clean match. The amount sits inside the customer's limit. Three green lights. So how much comfort should they give you?
Less than you'd hope. Confirmation of Payee is one of the few controls aimed squarely at scams, and scammers get past it with a story. "I can't use my own account right now, send it to my friend's; he can release the money today." The victim types the friend's real name, the name matches, and the payment goes. In the GCC, it's even easier: most banks show the payer a half-masked beneficiary name, and a few visible letters are easy to explain away when someone on the phone tells you what you'll see. The check worked exactly as designed. It confirmed a name the criminal had already supplied.
The limit tells you even less. A limit often reflects the bank's risk appetite, not how this customer behaves, and customers go past theirs anyway.
Now add context: the amount dwarfs what this customer usually sends, the recipient is new to them, and the payment lands right after an odd run of account changes. None of that trips a hard rule. Together, I would want someone to look before the money leaves.
That is what anomaly detection is for. It asks what is different from the expected pattern?, and every technique you own can answer that: rules, scenarios, models, network and content analysis. In my experience, it's how most fraud gets caught: something looked different from the customer it claimed to be. That makes it the closest thing we have to a universal tool, the Swiss Army knife of fraud prevention. Just don't mistake it for bulletproof.
The value of an anomaly is what it adds to the evidence you already have.
One detection concept across many techniques
Anomaly detection is not a technique. It is a concept that runs across your techniques, describing the search for departures from an expected pattern. In a taxonomy organized by technique, anomaly belongs at the pattern layer: it says what you are looking for, and the implementation says how you find it.
The statistical methods people usually mean by "anomaly detection" are one analytical implementation of that idea. 1
| Anomalous pattern | Possible implementation |
|---|---|
| Amount outside the customer's usual range | A rule comparing the amount with stored customer-profile values |
| Unusual order of account changes and payments | A scenario evaluating the event sequence against an expected journey |
| Activity unlike that of comparable customers | A model assessing distance from, or density within, a peer group |
| Unusual relationships between accounts | Network analysis examining relationship structure |
| Invoice details inconsistent with previous documents | Content analysis comparing extracted information |
The anomaly says what makes the activity noteworthy; the technique says how the evidence gets produced. That doesn't make every rule an anomaly detector. A payment above a contractual limit just breaches policy, and a known fraud script can be suspicious without being rare. If you want to call something anomalous, name the expectation it departs from. No expectation, no anomaly.
Normal compared with what?
Anomaly detection is only as good as its baseline. That baseline might be a simple statistical range or a pattern a model learned; it does not need an elaborate AI system behind it. 1
So the real question is: normal compared with what?
The reference can sit at three levels: the whole population, a peer group of similar customers, or the customer's own history. A large corporate payment can look extreme across all customers yet ordinary among comparable businesses, and ordinary for comparable businesses yet extreme for this one. Which level you end up with mostly depends on technology. Population thresholds are easy to set; a profile for every customer, updated as they transact, takes a platform built for it. Where you have the choice, go to the individual. Peer groups and population baselines are the fallback for customers with too little history, not the default.

The unit matters too: a single amount can be ordinary while a burst of similar payments is not. Clustering helps build peer groups, but on its own it isn't fraud detection. It becomes evidence only when you read a departure against those groups.
Take the GCC, where I've worked for over thirteen years. Rent is often paid in large lumps, quarterly, twice a year, sometimes a full year up front. School fees arrive the same way. Every summer, a large share of residents book flights and hotels within a few weeks, in a merchant category most fraud teams already treat as high risk. Measured against a 30- or 90-day baseline, each one is a textbook anomaly: many times the customer's usual amount, to a payee they haven't paid in months. Nothing about it is suspicious. The customer does the same thing every year.
Ramadan stretches "usual" further, and it starts before the month itself. In the run-up, households stock up, buy gifts, and settle bills, so the number of transactions in a short window rises sharply, which is exactly what velocity rules count. During the month, spending moves into the late evening and night, and charity and family transfers rise. Because Ramadan follows the lunar calendar, it starts about eleven days earlier each year, so a baseline that learned last year's March as normal is now looking at the wrong month.

The lesson I'd take from this: if your baseline window is shorter than the customer's payment cycle, you will keep flagging the normal behavior.
The amount alone never settles it, and neither does paying at 3 AM or logging in from a new device. Even a "new device" flag is only as reliable as your device recognition, because browser and device characteristics change on their own 2; my article on device fingerprinting digs into that.
Baselines also rot. A customer changes jobs and their old "normal" starts producing alerts. Absorb recent activity too eagerly, though, and you quietly normalize a slow fraudulent drift. The update policy matters as much as the method.
Putting the signals together
Once something looks unusual, ask what is driving it: what changed, what was it compared with, and what would move the assessment?
Start with the distinction that changes the whole investigation. In an account takeover, the attacker operates the account. In an authorized push payment scam, the real customer makes the payment while a criminal talks them through it, so a clean authentication can sit right alongside the fraud. A new-device signal can crack a takeover case open and add nothing when a scam victim is on their usual phone. There, the payment context, the recipient, and the recent shape of the customer's journey tell you more. Behavioral signals help too, as BioCatch's published case material shows, though I wouldn't read one vendor's headline number as a general promise. 3
When clients ask me for AI, they usually aren't asking for a model. They want a rule, or a set of rules, that can weigh many parameters at once and choose the right response. Most of them can already build anomaly rules: amount against the customer's average, a first payment to a new beneficiary. And they're right to sense where those rules run out. A person can hold two or three indicators in their head as one "pattern." We round thresholds to tidy numbers, and we split customers into the two or three most obvious groups. Analytics can combine far more signals and put the boundaries where the data says, not where a round number feels comfortable. That's the honest case for a model: it does the combining a person can't. It doesn't make the rules you already have wrong. I made the wider version of that argument in Before You Ask for AI, Check Your Pantry.
I go deeper on combining techniques in Fighting Fraud with Composite AI. One warning: "large payment" and "above the customer's usual amount" can be the same event wearing two hats (patterns). Count them twice, and you might be steering the final decisioning to decline without sufficient backing.
Testing what a flag adds
An anomaly flag has one job: improve the decision. So test it inside the logic you already run, not by how many alerts it raises on its own. Run the old and new versions over a later stretch of data with confirmed outcomes, and compare fraud captured, legitimate customers challenged, and analyst workload together. Look at the fraud you missed, not just the alerts you saved. A quieter queue is an operational result, not proof of better detection, and headline accuracy hides the difference, as my article on the confusion matrix explains.
Most anomaly flags start out noisy, and much of that noise is innocent: the rent, the school fees, the Ramadan run-up. Raising the threshold for everyone doesn't fix that. Looking for the good patterns as deliberately as the bad ones does. Most rule libraries are written only to find fraud, and very few have anything built to recognize the genuine customer: the payment that lands on the same date as last year's, to the same landlord; the beneficiary added weeks before it was first paid; the large payment to a verified institution, a school or a utility, rather than to a personal account. Those are departures from the usual too, but they point toward the real customer. None of them is proof, any more than a single fraud anomaly is. They're evidence on the other side of the scale, and a setup that can see them can let the payment through without loosening the rule for everyone else.
That's why, on every implementation I work on, we bring in as many fields from the payment message as is practical, including ones the business team never asked for. They're rarely needed on day one, but they're what lets you describe a good pattern precisely enough to trust it.
Turning evidence into action
An anomaly is a reason to look, not a verdict. The response should follow the strength and relevance of the combined evidence: continued monitoring, a warning aimed at the specific scam you suspect, a review, or a hold. A generic step-up authentication often tells you only that someone holds an authentication factor, not why the payment is being made, which is usually the question that matters.
Design the response next to the detection logic: what evidence justifies intervening, who reviews it, and whether they can act before the payment completes. There are more options than a blunt yes or no, as I argue in Approve or Decline - are these all our options?.
Most rules start in advisory mode: they raise alerts but don't stop payments. Moving a rule from advisory to blocking is a different decision from tuning it. Someone signs it off, usually after testing has built confidence in the false-positive rate and the other numbers. Gut feel still plays a part, but here teams work hard to back the call with figures, because a blocked payment can become a complaint that reaches executive management, and "it felt right" is no defense there.
So where should the line sit? I don't think there's a generic answer. It comes down to the scenario and the context you can gather. Sometimes a single anomaly is enough to stop a payment, when its size or its deviation from normal is large enough to carry high confidence on its own. More often, what matters is how ambiguous the anomaly is. "The customer is using a channel they've never used before" tells you very little. "An elderly customer is using mobile banking for the first time, and the beneficiary is a new international IBAN" starts to tell you a story.
That's the real test. More anomalies don't automatically mean more risk. Some build toward a fraud story. Others build toward an innocent one: the large transfer in the week the annual rent is due, from the customer's usual device, to the landlord they paid last year. Counting anomalies won't tell you which story you're looking at. Reading them together will.

What this means for you
Pick one fraud scenario and write down what each check actually establishes - the fact, not the feeling it gives you. For each anomaly, ask which story it supports, fraud or an innocent one, and check that its baseline window is longer than the customer's payment cycle. Then write down what the genuine version of that scenario looks like, and build checks that recognize it, not only the fraud. If you're implementing now, bring in more message fields than the business asked for; you'll need them once tuning starts. And when you tune, judge each change by fraud captured, needless interventions, and analyst workload together. And ensure that the change doesn't only quiet the queue, as you might not have improved detection but rather just stopped looking.
References & Further Reading
[1] scikit-learn - Novelty and Outlier Detection
Background on detecting departures from expected data patterns, anomaly scores, and thresholds.
[2] W3C - Mitigating Browser Fingerprinting in Web Specifications
Defines fingerprinting through observable characteristics and discusses factors affecting identification.
Vendor-reported case study illustrating behavioral analysis in social-engineering detection; its headline result is specific to that deployment.