Menu

System design · Case study 2 of 15

Design a Clickstream Data Platform

Design a platform that collects every page view and click from a website and mobile apps and turns it into reliable product analytics such as sessions, funnels and retention.

  • Advanced
  • 2 min read
  • Updated Oct 2026
On this page
  1. Approach
  2. Architecture
  3. Sessionisation
  4. Late and duplicate events
  5. Bots
  6. Privacy
  7. Observability
  8. Cost

Functional requirements

  • Collect events from web and mobile clients
  • Validate events against a schema and route invalid ones aside
  • Build sessions and daily user activity tables
  • Support funnel and retention analysis

Non-functional requirements

  • Collection endpoint highly available; clients never block on it
  • Raw events available within 5 minutes; modelled tables daily
  • Consent and privacy rules enforced
  • Handle bot traffic and duplicate events

Scale assumptions

  • About 2 billion events per day
  • Roughly 0.5 KB per event
  • Two years of history for analysis

Technologies

Collection API, Kafka, Schema registry, Spark (streaming and batch), Lakehouse table format

A lightweight collection service writes validated events to Kafka; a streaming job lands them in a partitioned lakehouse table; daily batch jobs build sessions and aggregates.

Approach

Clickstream is high-volume and messy. Focus on cheap, reliable collection, schema discipline, late mobile data and privacy, then on modelling sessions.

Architecture

  1. Clients batch events and send them to a collection endpoint; they never block the user interface.
  2. Collection service validates against the schema registry and writes valid events to Kafka, invalid ones to a dead-letter topic.
  3. Streaming landing appends raw events to a lakehouse table partitioned by event date.
  4. Daily batch deduplicates, filters bots, builds sessions and user-day tables.
  5. Analytics: funnels and retention queries run on the modelled tables.
Collection and landing are continuous; modelling is a daily, rerunnable batch.

Sessionisation

Order a user’s events by event time and start a new session after 30 minutes of inactivity (a common convention; make the threshold configurable). Implement with window functions: LAG of the timestamp, a flag when the gap exceeds the threshold, and a running sum of flags as the session number.

Late and duplicate events

Mobile clients send late. Reprocess the last few days’ partitions daily so late events land in the right sessions. Deduplicate on the client-generated event id.

Bots

Filter known bot user agents and implausible behaviour (for example thousands of events per minute from one id), and keep a flag rather than deleting so filters can be revised.

Privacy

Collect only consented events, pseudonymise user ids, restrict raw tables, and support deletion by user id across raw and modelled layers.

Observability

Events per client version, invalid-event rate, collection latency, Kafka lag and daily volume against baseline.

Cost

Columnar storage with compression, compaction of streaming files, lifecycle policies for raw data beyond the analysis window.

By Data Career Hub Editorial · Last reviewed Oct 2026

Search
Filter by type