Fraud Detection
Frequency: Reported
Capital One processes millions of transactions, most of which are legitimate. A small minority are fraudulent and need to be flagged for review.
You are given two datasets, transaction_values and transaction_dates. Parse and store the data, identify fraudulent outliers, and then discuss how the implementation should scale to multiple customers.
Initial data schemas
transaction_values contains one customer's transactions:
| Field | Type | Meaning |
|---|---|---|
transaction_id | integer | Unique transaction identifier |
transaction_type | string | Initially CREDIT; determines whether the signed value is positive or negative |
amount | double | USD magnitude with two decimal places |
transaction_values = [
[1, "CREDIT", 100.00],
[2, "CREDIT", 1000.00],
[3, "CREDIT", 25.15],
[100, "CREDIT", 15.21],
[245, "CREDIT", 72.30],
[311, "CREDIT", 25.19],
]transaction_dates contains the date of each transaction:
| Field | Type | Meaning |
|---|---|---|
transaction_id | integer | Unique transaction identifier |
date | string | Date in mmddyyyy format |
transaction_dates = [
[1, "03122022"],
[2, "04012022"],
[3, "04012022"],
[100, "04012022"],
[245, "04212022"],
[311, "04252022"],
]Level 1 - Store transactions
Create one or more methods that parse and store a single customer's transactions, including each transaction's date, using a data structure of your choice.
Discuss:
- the time and space complexity;
- why you chose the data structure; and
- whether the data should be indexed by date or transaction value.
Level 2 - Add debits
transaction_values can now contain DEBIT, meaning that money leaves the account. Debits may be treated as negative values. Modify the existing code to support this information.
transaction_values = [
[1, "CREDIT", 100.00],
[2, "CREDIT", 1000.00],
[3, "CREDIT", 25.15],
[100, "DEBIT", 15.21],
[245, "DEBIT", 72.30],
[311, "DEBIT", 25.19],
]Discuss:
- how the time and space complexity changed;
- the reason for any new data structure; and
- whether and how the solution should change for 10,000 entries.
Level 3 - Flag fraudulent transactions
One possible fraud rule compares a transaction with the average transaction for that client. A transaction is likely fraudulent when its value is greater than twice the mean for the same transaction type (CREDIT or DEBIT).
Identify and output fraudulent transactions for further investigation.
Example:
transaction_values = [
[1, "CREDIT", 10.00],
[2, "DEBIT", 10.00],
]
transaction_dates = [
[1, "03122022"],
[2, "03122022"],
]
expected = []Larger example:
transaction_values = [
[1, "CREDIT", 100.00],
[2, "CREDIT", 1000.00],
[3, "CREDIT", 25.15],
[100, "DEBIT", 15.21],
[245, "DEBIT", 72.30],
[311, "DEBIT", 25.19],
]
transaction_dates = [
[1, "03122022"],
[2, "04012022"],
[3, "04012022"],
[100, "04012022"],
[245, "04212022"],
[311, "04252022"],
]
expected = [
[2, "CREDIT", 1000.00, "04012022"],
]Again discuss complexity, data-structure choices, and changes needed at 10,000 entries.
Level 4 - Multiple clients
Extend the application across multiple clients. Discuss time and space complexity and the implications of one million customers averaging 50 transactions per month.
The transaction-value schema now includes client_id:
| Field | Type | Meaning |
|---|---|---|
client_id | integer | Unique client identifier |
transaction_id | integer | Unique transaction identifier |
transaction_type | string | CREDIT or DEBIT |
amount | double | USD magnitude with two decimal places |
transaction_values = [
[1, 1, "CREDIT", 100.00],
[2, 2, "CREDIT", 1000.00],
[1, 3, "CREDIT", 25.15],
[1, 100, "DEBIT", 15.21],
[1, 245, "DEBIT", 72.30],
[1, 311, "DEBIT", 25.19],
]
transaction_dates = [
[1, "03122022"],
[2, "04012022"],
[3, "04012022"],
[100, "04012022"],
[245, "04212022"],
[311, "04252022"],
]
expected = []Reported answer
The archive includes a candidate's Python implementation. It is preserved as reported, including possible bugs and unfinished choices: reported solution.