
How (not) to train the fraud detection model! (Part 2)
Milo improves his fraud model with transaction sequences, non-financial events, behavioral profiles and third-party data, discovering the work behind useful features.
Milo felt stuck again, and since he last spoke with fraud department colleagues, he had improved the model; he decided to try his luck again and talk to them once more. When he met them, he told them about his recent model improvements and that he couldn't see any other way to improve it.
Leveraging historical memory (sequence of events & entity profiles)
After listening to Milo, one of the fraud analysts noticed that Milo's model assessed transactions as individual and independent; it used no memory to look at what was happening before the transaction. Analysts explained that some of their business rules look for sequences of transactions, and that multiple transactions within a short time are simple examples of high-risk events.

This was an entirely new way of looking at the problem. So he extended his ABT in his next revision cycle and added additional columns: "number of transactions within the last four hours," "time since the last transaction," and a few more. He felt that this might be a game-changer. But to calculate these details, he had to look at every transaction in his ABT individually and consider the transactions before that moment within the predefined time window. This wasn't a trivial task for him as a data scientist, so he asked his data management colleagues in IT for help and advice on ETL building. With their help, he prepared the columns. As he expected, the results were really good. In the next few iterations, he experimented with different time windows to find the best fit for the model.
After a few days, Milo felt that he had plateaued and that further fine-tuning the time windows and trying different sequences would bring no significant gains. He enthusiastically went to see the fraud risk colleagues again.
Leveraging non-financial transactions (sequence of events & entity profiles)
During this discussion, the fraud analyst realized that Milo might have left early the last time as he forgot to tell him about non-financial transactions. How extremely beneficial those might be in the same context of using them as part of sequences. He also mentioned a couple of examples: a change of mobile number shortly before a funds transfer or activation of the new channel before using it for a funds transfer. Both patterns significantly increase the fraud risk. Milo realized he wasn't done with the transaction sequences yet, but he also remembered he needed help from his data management colleagues to get the code for calculating these extra columns.
Another problem lurked around the corner: getting the non-financial transactions. So Milo went back to his data management colleagues to figure out how big a problem he was facing now. He briefly explained his high-level objective of building a predictive model for fraud detection and how he wanted to include non-financial transactions in the mix. Although his colleagues were interested and the meeting ran far beyond the scheduled time, he shared only high-level information without precise details, as his colleagues from the fraud risk team guided him.

The outcome of the discussion was not a complete disaster, but it wasn't a clear victory either. For example, some non-financial transactions (like a change of mobile number, change of address, or request to issue a new checkbook) were available in the DWH in a specific table. But other events like "successful login" or "failed login" weren't available, since the channel itself managed those events. Colleagues also told him they could potentially bring those events (and others) into the DWH, but until then, there was no clear use case for that data.
Still, the hard part would be transforming the non-financial events into a meaningful flag like "was there a change of mobile number within 24h" or "was there a registration of a new trusted device within 24h". But the colleagues from the data management team came to the rescue and helped Milo prepare the code to derive the required details.
Another round of training and validation brought positive news, and the model's accuracy improved again. Looking back at all the different data that ABT was working with, Milo was quite impressed. Yet he wanted to know whether they could add more details that might become better fraud indicators than the ones already identified. So, without any hesitation, he went back to his fraud-risk colleagues to brainstorm potential new ideas.
Anomaly detection & behavioral characteristics
After sharing his latest adjustments and their contribution to accuracy, they decided to discuss the last few fraud incidents together and analyze the patterns behind them to see whether they could identify a new approach.
While reviewing the details, they saw that most frauds in this typology are linked to social engineering techniques and often end up as account takeover frauds. Fraud analysts mentioned that they had deployed rules to look for unusual customer behavior. They checked whether customers' present amounts and volumes of transactions were not uncommon (too high), whether the channel was used in alignment with previous observations, and some other patterns. To Milo, this could improve accuracy and reduce false positives, as the model could spot abnormal or anomalous customer behavior more accurately. 
"So what would it mean to build such profiles?" Milo asked himself. Thinking about it, he realized that he might need additional data, even before 6 months, if he wanted to build a profile for all transactions within his ABT. For example, suppose he wanted to calculate the "maximum of the cumulative daily amount of debits for the given customer within the last month" for today's transaction. In that case, he would need the previous month's transactions, calculate the cumulative daily debit amount for each customer for each day, and then select the maximum across those 30 days for each customer. For transactions from 6 months back, he would need data from 1 month prior to the 7th month. Extending the data based on the depth of the historical window was the first complication.
Each characteristic is linked to a particular entity - most commonly a customer. So, for example, we can calculate the number of transactions the customer has done in 24h; similarly, we can calculate the statistics on the account level - the number of transactions in 24h for each account independently or each channel alone. We can also calculate other metrics - minimum, average, percentiles, etc. For example, we can sum amounts or counts daily, weekly, and monthly; we can consider any transaction or be very narrow and count only debit transactions performed via a branch to a particular country, etc. Other entities could be an employee or product like a credit card, debit card, checkbook, etc.
The options were limitless, and only now did Milo realize how important it is to understand fraud schemes to design and calculate new columns that would eventually become part of the model and help distinguish fraud from genuine transactions.
Enrichment via 3rd party data
Milo knew his visit to fraud risk colleagues was just the beginning, and to improve the model further, he needed to gather as many insights as possible.
When he was reading about different detection techniques - especially the ones related to payment fraud perpetrated through e-channels - he found an article that mentioned a new term: device fingerprinting. Milo was surprised that some scripts could capture both static and behavioral characteristics. While the web visitor browsed the pages, the script captured static details like device type (PC, iOS, Mac, Android), device language settings, screen resolution, firmware version, timezone, IP, browser details, and more. It was also possible to capture specific behavioral characteristics like typing speed, mouse cursor movement, visit duration on each page, and others.
If these details were available and could be added to the ABT, they would further improve accuracy and reduce false positives. But Milo would need to discuss this with a broader audience within the bank, as integrating external 3rd-party data would require a project of its own.
Milo closed the browser with mixed feelings, realizing his journey to further enhance the model would never end. Even though he felt he had reached the end many iterations back, it was clear that the actual end was nowhere near his reach.
Did Milo successfully deploy his model into production? Did he make any mistake that will cause trouble while Milo tries to operationalize the model?
To Be Continued …
Disclaimer: The story, all names, characters, and incidents portrayed in this article are fictitious. No identification with actual persons (living or deceased), places, buildings, and products is intended or should be inferred.
