Course · Cloud
AWS
AWS for Data Engineers: S3, IAM, Lambda, Glue, Athena, Redshift, EMR, Kinesis, Step Functions, Lake Formation and how they fit into one data platform.
- Lessons
- 2
- Interview questions
- 0
- Projects & case studies
- 2
- Reading time
- ~1 h
About this course
Amazon Web Services is the most common cloud in Data Engineer job descriptions. An AWS data platform is usually a lake in Amazon S3, described by the AWS Glue Data Catalog, loaded by streaming and batch ingestion, transformed with Spark on Glue or EMR, queried with Athena or Redshift, orchestrated with Step Functions or Airflow, governed with IAM and Lake Formation, and watched with CloudWatch.
Start with the map of the AWS data stack, then S3 and IAM, because every other service depends on them. After that, follow the lessons in order or jump to the service your team uses.
Your progress
Saved in this browser onlyPractise
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
System design case studies
Resources
Related courses
- Apache SparkUnderstand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.
- KafkaKafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.
- AirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.
- SnowflakeSnowflake is a cloud data warehouse that separates storage from compute. Learn virtual warehouses, micro-partitions, pruning, caching and cost control.