Menu

Advanced project · Project 8 of 8

Fraud Detection Data Pipeline

Build the data side of a fraud-detection system: compute per-card behavioural features from a transaction stream, flag suspicious transactions with transparent rules, and maintain a feature table that a model could use.

  • Advanced
  • Kafka · Spark Structured Streaming · Python · Delta Lake or PostgreSQL
  • 2 min read
  • Updated Oct 2026

Requirements

  • Stream synthetic card transactions through Kafka
  • Compute rolling features per card (count and amount in the last 10 minutes and 24 hours, distinct merchants)
  • Flag transactions with explainable rules
  • Write flags to a review queue and features to a table
  • Keep raw transactions for backtesting rules

Technology stack

Kafka, Spark Structured Streaming, Python, Delta Lake or PostgreSQL

Dataset

Generate synthetic transactions with injected fraud-like patterns, or use a public synthetic fraud dataset whose licence permits reuse. Do not use real card data.

Business context

Fraud teams need suspicious transactions surfaced quickly, with a reason they can act on. Data engineers build the features and the plumbing; this project focuses on that part and keeps the detection logic transparent.

Architecture

  1. Transactions arrive on a Kafka topic.
  2. Spark computes rolling features per card.
  3. Features join each transaction; rules evaluate them.
  4. Flagged transactions go to a review table with the triggering rule.
  5. Raw transactions and features are stored for backtesting and modelling.
Features are computed once and used both for live rules and for later model training.

When you report detection results, report them on your synthetic data and describe how the patterns were generated. Results on synthetic data do not show real-world performance.

By Data Career Hub Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type