Menu

Python course · Lesson 3 of 5

Python Data Structures for Data Engineering Interviews

Choose between list, tuple, set, dict, deque and Counter by their operations and costs, with the patterns that come up in data engineering coding interviews.

  • Beginner
  • 3 min read
  • Updated Oct 2026
On this page
  1. Membership: use a set
  2. Deduplicate while keeping order
  3. Grouping and counting
  4. Queues and sliding windows
  5. Tuples as composite keys
  6. Common mistakes
  7. Interview relevance
  8. Key takeaway

Choosing the right built-in structure is often the whole answer to a coding question. Know what each one is for and what its common operations cost.

Structure Ordered Mutable Duplicates Lookup by value Typical use
list Yes Yes Yes O(n) scan Ordered records, stacks
tuple Yes No Yes O(n) scan Fixed records, dict keys
set No Yes No O(1) average Membership, deduplication
dict Insertion order Yes Unique keys O(1) average by key Lookups, grouping, counting
collections.deque Yes Yes Yes O(n) scan Queues: fast appends/pops at both ends

Membership: use a set

seen_ids = {101, 102, 103}
print(102 in seen_ids, 999 in seen_ids)
True False

Checking x in some_list scans the list, which gets slow inside a loop over a large input. Converting to a set first turns repeated membership checks from O(n) into O(1) on average.

Deduplicate while keeping order

events = ["b", "a", "b", "c", "a"]
print(list(dict.fromkeys(events)))
['b', 'a', 'c']

Dicts preserve insertion order (guaranteed since Python 3.7), so dict.fromkeys keeps the first occurrence of each value. list(set(events)) also deduplicates but loses order.

Grouping and counting

from collections import Counter, defaultdict

orders = [("asha", 50), ("ben", 20), ("asha", 70)]

totals = defaultdict(int)
for customer, amount in orders:
    totals[customer] += amount
print(dict(totals))

print(Counter(customer for customer, _ in orders).most_common(1))
{'asha': 120, 'ben': 20}
[('asha', 2)]

These are the in-memory equivalents of GROUP BY with SUM and COUNT.

Queues and sliding windows

from collections import deque

window = deque(maxlen=3)
for value in [1, 2, 3, 4, 5]:
    window.append(value)
print(list(window), sum(window) / len(window))
[3, 4, 5] 4.0

deque(maxlen=n) drops old items automatically, which is handy for moving averages over a stream.

Tuples as composite keys

Tuples are immutable, and therefore hashable when their elements are, so they can be dict keys or set members:

daily = {("2026-10-01", "north"): 150, ("2026-10-01", "south"): 230}
print(daily[("2026-10-01", "south")])
230

Common mistakes

  1. Membership checks against a list inside a loop (quadratic time).
  2. Using a list as a queue with pop(0), which shifts every element. Use deque.popleft().
  3. Trying to use a list as a dict key (lists are not hashable).
  4. Assuming sets keep order.

Interview relevance

Many “process these records” questions reduce to: group with a dict, deduplicate with a set or dict.fromkeys, count with Counter, and stream with a deque. State the time complexity of your choice.

Key takeaway

Pick the structure by the operation you do most: sets for membership, dicts for lookup and grouping, deques for queues, tuples for fixed records and keys.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Examples run on Python 3.12

Progress is saved in this browser only. No account needed.

Search
Filter by type