Menu

Set a start date to see a calendar date on each day and highlight today. Saved in this browser only.

Phase 1: Foundations

Week 1SQL joins/windows solid + 15 DSA done

  1. Day 1SQL + DSA focus3 h
  2. Day 2Big Data Engineering3 h
  3. Day 3Cloud & Warehousing3 h
  4. Day 4Streaming Systems3 h
  5. Day 5System Design Deep-Dive6 h
  6. Day 6Data Modeling & SQL6 h
  7. Day 7Mock Interview Day3 h

Week 2Advanced SQL + 30 DSA + Spark core

  1. Day 8Big Data Engineering3 h
  2. Day 9Cloud & Warehousing3 h
  3. Day 10Streaming Systems3 h
  4. Day 11System Design Deep-Dive3 h
    • SQLIN, BETWEEN, LIKE patterns
    • DSAMerge Intervals
    • PySparkreduceByKey vs groupByKey
    • SnowflakeVirtual warehouse sizing
    • Kafkaacks (0/1/all)
    • AirflowCustom operators
    • AWSIAM users/groups/roles
    • System DesignDesign a metrics / KPI platform
    • AlsoRevise notes / flashcards
  5. Day 12Data Modeling & SQL6 h
    • SQLNULL handling with IS NULL / COALESCE
    • DSAInsert Interval
    • PySparkmapPartitions
    • SnowflakeScaling up vs scaling out
    • KafkaIdempotent producer
    • AirflowHooks
    • AWSPolicies (identity/resource)
    • System DesignDesign an ELT pipeline with dbt
    • AlsoApply to 3 target companies
  6. Day 13SQL + DSA focus6 h
  7. Day 14Mock Interview Day3 h
    • SQLString functions (CONCAT, SUBSTRING, TRIM)
    • DSARotate Image
    • PySparkCheckpointing
    • SnowflakeAuto-suspend & auto-resume
    • KafkaCompression (snappy/lz4/zstd)
    • AirflowKubernetesPodOperator
    • AWSAssume role & STS
    • System DesignDesign a data quality framework
    • AlsoFull mock interview + review

Week 3PySpark performance + Snowflake basics

  1. Day 15Cloud & Warehousing3 h
    • SQLDate functions basics
    • DSASpiral Matrix
    • PySparkAccumulators
    • SnowflakeWarehouse concurrency
    • KafkaPartitioner strategies
    • AirflowBranching (BranchPythonOperator)
    • AWSCross-account access
    • System DesignDesign a data observability system
    • AlsoApply to 3 target companies
  2. Day 16Streaming Systems3 h
    • SQLCAST and data type conversion
    • DSASet Matrix Zeroes
    • PySparkBroadcast variables
    • SnowflakeQuery queuing
    • KafkaRetries & delivery
    • AirflowTrigger rules
    • AWSService-linked roles
    • System DesignDesign a data catalog & lineage system
    • AlsoRevise notes / flashcards
  3. Day 17System Design Deep-Dive3 h
  4. Day 18Data Modeling & SQL3 h
  5. Day 19SQL + DSA focus6 h
  6. Day 20Big Data Engineering6 h
  7. Day 21Mock Interview Day3 h

Week 41 end-to-end project + 50 DSA

  1. Day 22Streaming Systems3 h
  2. Day 23System Design Deep-Dive3 h
  3. Day 24Data Modeling & SQL3 h
  4. Day 25SQL + DSA focus3 h
    • SQLAnti-join (NOT EXISTS / LEFT JOIN NULL)
    • DSASort Colors
    • PySparkgroupBy & aggregations
    • SnowflakeSearch optimization service
    • KafkaDeserialization
    • AirflowTimetables
    • AWSError handling & DLQ
    • System DesignDesign a cost-optimized warehouse strategy
    • AlsoRevise notes / flashcards
  5. Day 26Big Data Engineering6 h
  6. Day 27Cloud & Warehousing6 h
  7. Day 28Mock Interview Day3 h

Week 5Kafka + streaming fundamentals

  1. Day 29System Design Deep-Dive3 h
  2. Day 30Data Modeling & SQL3 h

    Month 1: Core SQL+DSA+Spark mastered; 1 project; start applying Target: 60% prep • first OAs & recruiter calls.

Phase 2: Advanced + Projects

Week 5Kafka + streaming fundamentals

  1. Day 31SQL + DSA focus3 h
  2. Day 32Big Data Engineering3 h
    • SQLMultiple chained CTEs
    • DSAMaximum Average Subarray I
    • PySparkReading/writing JSON & CSV
    • SnowflakeTask fundamentals
    • KafkaPartition reassignment
    • AirflowParams
    • AWSGlue bookmarks
    • System DesignDesign a time-series metrics store
    • AlsoRevise notes / flashcards
  3. Day 33Cloud & Warehousing6 h
    • SQLConditional aggregation
    • DSAValid Parentheses
    • PySparkHandling nested/complex types
    • SnowflakeScheduled tasks (CRON)
    • KafkaLeader & followers
    • AirflowTask state lifecycle
    • AWSGlue triggers/workflows
    • System DesignDesign a search indexing pipeline
    • AlsoApply to 3 target companies
  4. Day 34Streaming Systems6 h
    • SQLPIVOT rows to columns
    • DSAMin Stack
    • PySparkexplode & posexplode
    • SnowflakeTask trees / DAGs
    • KafkaIn-Sync Replicas (ISR)
    • AirflowSequentialExecutor
    • AWSGlue partitions
    • System DesignDesign a notification/alerting pipeline
    • AlsoRevise notes / flashcards
  5. Day 35Mock Interview Day3 h
    • SQLUNPIVOT columns to rows
    • DSAEvaluate Reverse Polish Notation
    • PySparkpivot in Spark
    • SnowflakeServerless tasks
    • Kafkamin.insync.replicas
    • AirflowLocalExecutor
    • AWSPySpark on Glue
    • System DesignDesign an order-events processing system
    • AlsoFull mock interview + review

Week 6Airflow + AWS Glue/Redshift

  1. Day 36Data Modeling & SQL3 h
    • SQLGROUPING SETS
    • DSAGenerate Parentheses
    • PySparkTemp views & global views
    • SnowflakeTask + Stream pattern
    • KafkaUnclean leader election
    • AirflowCeleryExecutor
    • AWSGlue cost tuning
    • System DesignDesign a payment events pipeline (exactly-once)
    • AlsoApply to 3 target companies
  2. Day 37SQL + DSA focus3 h
    • SQLROLLUP and CUBE
    • DSADaily Temperatures
    • PySparkSQL vs DataFrame performance
    • SnowflakeTask error handling
    • KafkaHigh watermark
    • AirflowKubernetesExecutor
    • AWSAthena basics
    • System DesignDesign a ride-hailing surge pricing pipeline
    • AlsoRevise notes / flashcards
  3. Day 38Big Data Engineering3 h
  4. Day 39Cloud & Warehousing3 h
  5. Day 40Streaming Systems6 h
  6. Day 41System Design Deep-Dive6 h
  7. Day 42Mock Interview Day3 h

Week 7Data modeling + 80 DSA

  1. Day 43SQL + DSA focus3 h
  2. Day 44Big Data Engineering3 h
  3. Day 45Cloud & Warehousing3 h
    • SQLRunning totals with SUM OVER
    • DSABinary Search
    • PySparkPredicate & projection pushdown
    • SnowflakeTime Travel basics
    • KafkaIdempotence end-to-end
    • AirflowAirflow REST API
    • AWSRedshift architecture
    • System DesignDesign a geospatial analytics pipeline
    • AlsoApply to 3 target companies
  4. Day 46Streaming Systems3 h
    • SQLMoving average with frame clause
    • DSASearch a 2D Matrix
    • PySparkCaching strategy
    • SnowflakeAT / BEFORE clauses
    • KafkaRead-process-write pattern
    • AirflowConnections & secrets backend
    • AWSDistribution styles
    • System DesignDesign an ad-bidding analytics pipeline
    • AlsoRevise notes / flashcards
  5. Day 47System Design Deep-Dive6 h
    • SQLPercent of total
    • DSAKoko Eating Bananas
    • PySparkBucketing in Spark
    • SnowflakeUNDROP
    • KafkaConnect framework
    • AirflowIdempotent tasks
    • AWSSort keys
    • System DesignDesign a backfill & late-data handling system
    • AlsoRevise notes / flashcards
  6. Day 48Data Modeling & SQL6 h
  7. Day 49Mock Interview Day3 h

Week 82nd project + system design start

  1. Day 50Big Data Engineering3 h
    • SQLGenerate date series / calendar table
    • DSATime Based Key-Value Store
    • PySparkSpill to disk diagnosis
    • SnowflakeCloning with Time Travel
    • KafkaConverters & SMTs
    • AirflowCI/CD for DAGs
    • AWSConcurrency scaling
    • System DesignDesign an LLM/RAG data ingestion pipeline
    • AlsoRevise notes / flashcards
  2. Day 51Cloud & Warehousing3 h
  3. Day 52Streaming Systems3 h
  4. Day 53System Design Deep-Dive3 h
  5. Day 54Data Modeling & SQL6 h
  6. Day 55SQL + DSA focus6 h
  7. Day 56Mock Interview Day3 h

Week 95 system design scenarios + mocks

  1. Day 57Cloud & Warehousing3 h
  2. Day 58Streaming Systems3 h
  3. Day 59System Design Deep-Dive3 h
    • SQLMedian & percentile (PERCENTILE_CONT)
    • DSAAdd Two Numbers
    • PySparkAvoiding wide transformations
    • SnowflakeSecure views
    • KafkaState stores
    • AirflowDAG params & templating (Jinja)
    • AWSBootstrap actions
    • System DesignDesign a Customer 360 platform
    • AlsoRevise notes / flashcards
  4. Day 60Data Modeling & SQL3 h

    Month 2: Cloud+Streaming+Modeling done; 2 projects; active interviews Target: 85% prep • multiple onsites in pipeline.

Phase 3: Interview Mastery

Week 95 system design scenarios + mocks

  1. Day 61SQL + DSA focus6 h
    • SQLDeduplicate keeping latest record
    • DSALRU Cache
    • PySparkFile size optimization
    • SnowflakeWarehouse right-sizing
    • KafkaProcessor API
    • AirflowBashOperator
    • AWSCost optimization on EMR
    • System DesignDesign a recommendation data pipeline
    • AlsoRevise notes / flashcards
  2. Day 62Big Data Engineering6 h
    • SQLYear-over-year growth
    • DSAMerge k Sorted Lists
    • PySparkSmall files problem
    • SnowflakeAuto-suspend tuning
    • KafkaExactly-once in Streams
    • AirflowPythonOperator
    • AWSKinesis Data Streams
    • System DesignDesign a feature store
    • AlsoRevise notes / flashcards
  3. Day 63Mock Interview Day3 h
    • SQLMonth-over-month change
    • DSAReverse Nodes in k-Group
    • PySparkColumn pruning
    • SnowflakeQuery result caching
    • KafkaSchema Registry basics
    • AirflowCustom operators
    • AWSShards & throughput
    • System DesignDesign a metrics / KPI platform
    • AlsoFull mock interview + review

Week 10110 DSA + revise weak areas

  1. Day 64Streaming Systems3 h
    • SQLRolling 7/30 day metrics
    • DSAInvert Binary Tree
    • PySparkCaching vs recompute trade-off
    • SnowflakeAvoiding spilling
    • KafkaAvro/Protobuf/JSON schemas
    • AirflowHooks
    • AWSKinesis Firehose
    • System DesignDesign an ELT pipeline with dbt
    • AlsoRevise notes / flashcards
  2. Day 65System Design Deep-Dive3 h
  3. Day 66Data Modeling & SQL3 h
  4. Day 67SQL + DSA focus3 h
    • SQLSlowly changing dimension queries
    • DSABalanced Binary Tree
    • PySparkCost-based optimization
    • SnowflakeMaterialized view costs
    • KafkaSerializers/deserializers
    • AirflowBranching (BranchPythonOperator)
    • AWSKCL/KPL
    • System DesignDesign a data observability system
    • AlsoRevise notes / flashcards
  5. Day 68Big Data Engineering6 h
    • SQLDetecting consecutive streaks
    • DSASame Tree
    • PySparkStructured Streaming model
    • SnowflakeSecure Data Sharing
    • KafkaSchema references
    • AirflowTrigger rules
    • AWSKinesis Data Analytics
    • System DesignDesign a data catalog & lineage system
    • AlsoRevise notes / flashcards
  6. Day 69Cloud & Warehousing6 h
    • SQLPivoting dynamic columns
    • DSASubtree of Another Tree
    • PySparkInput sources (Kafka/files)
    • SnowflakeReader accounts
    • KafkaCluster sizing
    • AirflowSensor basics
    • AWSState machines
    • System DesignDesign a GDPR/PII compliant pipeline
    • AlsoApply to 3 target companies
  7. Day 70Mock Interview Day3 h
    • SQLConditional window frames
    • DSABinary Tree Level Order Traversal
    • PySparkOutput sinks & modes
    • SnowflakeData Marketplace
    • KafkaMonitoring & JMX metrics
    • AirflowPoke vs reschedule mode
    • AWSStandard vs Express
    • System DesignDesign a clickstream analytics pipeline
    • AlsoFull mock interview + review

Week 11Full mock loops + behavioral prep

  1. Day 71System Design Deep-Dive3 h
    • SQLEXCEPT / INTERSECT set operations
    • DSABinary Tree Right Side View
    • PySparkTriggers & micro-batch
    • SnowflakeListings & exchanges
    • KafkaThroughput tuning
    • AirflowExternalTaskSensor
    • AWSMap & Parallel states
    • System DesignDesign an IoT sensor data pipeline
    • AlsoRevise notes / flashcards
  2. Day 72Data Modeling & SQL3 h
    • SQLLATERAL / CROSS APPLY joins
    • DSACount Good Nodes in Binary Tree
    • PySparkWatermarking
    • SnowflakeCross-region/cloud sharing
    • KafkaQuotas
    • AirflowFileSensor
    • AWSError handling & retries
    • System DesignDesign a log ingestion & search platform
    • AlsoApply to 3 target companies
  3. Day 73SQL + DSA focus3 h
  4. Day 74Big Data Engineering3 h
    • SQLArray & nested data handling
    • DSABinary Tree Maximum Path Sum
    • PySparkStateful aggregations
    • SnowflakeZero-copy cloning
    • KafkaMirror Maker 2
    • AirflowSmart sensors
    • AWSOrchestration patterns
    • System DesignDesign a slowly changing dimension framework
    • AlsoRevise notes / flashcards
  5. Day 75Cloud & Warehousing6 h
  6. Day 76Streaming Systems6 h
  7. Day 77Mock Interview Day3 h

Week 12Final revision + apply aggressively

  1. Day 78Data Modeling & SQL3 h
    • SQLIndex design (B-tree, composite)
    • DSAKth Smallest Element in a BST
    • PySparkExactly-once in streaming
    • SnowflakeFLATTEN function
    • KafkaTombstones & log compaction
    • AirflowData-aware scheduling (Datasets)
    • AWSDashboards
    • System DesignDesign a data mesh architecture
    • AlsoApply to 3 target companies
  2. Day 79SQL + DSA focus3 h
  3. Day 80Big Data Engineering3 h
  4. Day 81Cloud & Warehousing3 h
  5. Day 82Streaming Systems6 h
  6. Day 83System Design Deep-Dive6 h
  7. Day 84Mock Interview Day3 h

Week 13Offer negotiation & decision

  1. Day 85SQL + DSA focus3 h
  2. Day 86Big Data Engineering3 h
    • SQLJoin algorithms (hash, merge, nested loop)
    • DSAKth Largest Element in a Stream
    • PySparkMERGE / upsert
    • SnowflakeMetadata management
    • KafkaRetention policies
    • AirflowSequentialExecutor
    • AWSCross-account data sharing
    • System DesignDesign a notification/alerting pipeline
    • AlsoRevise notes / flashcards
  3. Day 87Cloud & Warehousing3 h
    • SQLStatistics & the optimizer
    • DSALast Stone Weight
    • PySparkSchema enforcement
    • SnowflakeMulti-cluster shared data
    • KafkaCompaction (log cleanup)
    • AirflowLocalExecutor
    • AWSVPC & networking basics
    • System DesignDesign an order-events processing system
    • AlsoApply to 3 target companies
  4. Day 88Streaming Systems3 h
    • SQLDeadlocks & locking
    • DSAK Closest Points to Origin
    • PySparkSchema evolution
    • SnowflakeSnowflake editions
    • KafkaTopic configuration
    • AirflowCeleryExecutor
    • AWSDynamoDB for DE
    • System DesignDesign a payment events pipeline (exactly-once)
    • AlsoRevise notes / flashcards
  5. Day 89System Design Deep-Dive6 h
    • SQLIsolation levels (ACID)
    • DSAKth Largest Element in an Array
    • PySparkOPTIMIZE & compaction
    • SnowflakePricing model (credits)
    • KafkaProducer API
    • AirflowKubernetesExecutor
    • AWSRDS/Aurora basics
    • System DesignDesign a ride-hailing surge pricing pipeline
    • AlsoRevise notes / flashcards
  6. Day 90Data Modeling & SQL6 h
    • SQLMVCC concepts
    • DSATask Scheduler
    • PySparkZ-Ordering
    • SnowflakeCaching layers (result/local/metadata)
    • Kafkaacks (0/1/all)
    • AirflowCeleryKubernetes hybrid
    • AWSSQS & SNS
    • System DesignDesign a social media feed analytics system
    • AlsoApply to 3 target companies

    Month 3: Interview-ready across stack; multiple offers; negotiate offers Target: 100% prep • offers in hand.

Search
Filter by type