Delta Lake courseLesson 2 of 6
Delta Lake course · Lesson 2 of 6
JSON vs Parquet for Analytics Pipelines
Why raw JSON is convenient to land but expensive to query, how Parquet fixes size, types and scan cost, and a practical pattern for converting between them.
On this page
APIs and applications produce JSON, so pipelines receive a lot of it. JSON is a fine landing format and a poor analytics format.
| JSON (newline-delimited) | Parquet | |
|---|---|---|
| Human-readable | Yes | No |
| Schema | Implicit; inferred on read | Explicit, stored in the file |
| Types | Strings, numbers, booleans; dates are strings | Rich types (dates, timestamps, decimals) |
| Size | Large: field names repeated on every row | Compact columnar compression |
| Read a few columns | Must parse every row fully | Reads only needed columns |
| Nested data | Natural | Supported (structs, arrays, maps) |
The cost of querying JSON
In a PySpark 4.2 test, the same 200,000 rows took 21,260 KB as JSON and 2,382 KB as Parquet. Size is only part of it: every JSON query parses text and repeated field names, and schema inference requires an extra pass over the data. Inference also has side effects you might not expect; Spark’s inferred JSON schema lists fields in alphabetical order rather than the order they appear in the data.
A practical pattern
- Land raw JSON unchanged (bronze), partitioned by arrival date. It preserves exactly what the source sent.
- Parse with an explicit schema, cast types (timestamps, decimals), and capture malformed records.
- Write Parquet or a Parquet-based table format (silver) for all downstream use.
from pyspark.sql import SparkSession
spark = SparkSession.builder.getOrCreate()
schema = "event_id STRING, user_id BIGINT, event_type STRING, event_ts TIMESTAMP, amount DECIMAL(12,2)"
(spark.read.schema(schema).json("raw/events/dt=2026-10-01/")
.write.mode("overwrite").parquet("silver/events/dt=2026-10-01/"))
When to keep JSON
- Raw landing and replay.
- Small configuration or reference files.
- Truly schemaless payloads you store as a single column for later parsing.
Common mistakes
- Building dashboards directly on raw JSON.
- Relying on schema inference in production.
- Storing timestamps as strings in the analytics layer.
Key takeaway
Land JSON, analyse Parquet. Convert once, with an explicit schema, as early as possible.
Progress is saved in this browser only. No account needed.