Skip to content
ByDataFix
  • Home
  • Contact
  • Privacy Policy
  • Terms & Conditions
Subscribe
Subscribe
Delta Lake

Liquid Clustering vs Z-ORDER vs Partitioning: Which to Choose

Three ways to lay out a Delta table so queries read less data — and the senior skill isn't picking the newest one, it's knowing when NOT to switch. Explained with a bookstore.

Read More
July 25, 2026
Delta Lake

How Delta Lake Does ACID Without a Database

No database server, no lock manager — just files and a log. Here's how two Spark jobs can write the same table at the exact same moment and never corrupt it.

Read More
July 25, 2026
Apache Spark

Spark Driver OOM: Why collect() and Broadcast Joins Crash Your Job

Your executors have 64 GB and the job still dies with OutOfMemoryError. The culprit is the one part of Spark everyone forgets about — the Driver. Here's why it crashes and how to size it, in plain English.

Read More
July 25, 2026
Delta Lake

Delta Table Internals Explained: Files, Transaction Log, ACID & Time Travel

Follow one Delta table from CREATE to time travel and watch exactly what happens on disk — the Parquet files, the transaction log, and how that log quietly powers ACID and versioning.

Read More
July 25, 2026
Databricks

Medallion Architecture Explained Simply: Bronze, Silver & Gold

Raw data goes in one end, business-ready data comes out the other — refined in three clear stages. Here's the Bronze, Silver, and Gold pattern (plus the semantic layer), traced through a real production pipeline.

Read More
July 25, 2026
Apache Spark

repartition vs coalesce in Spark: Which to Use and When

Both change the number of partitions — but one does a full shuffle and one doesn't, and that single difference decides your performance. Plus the coalesce(1) trap everyone falls into.

Read More
July 24, 2026
Apache Spark

Spark Architecture Explained Simply: Driver, Executors, Jobs, Stages & Tasks

The words all sound the same and that's why it's confusing. One kitchen analogy makes the whole thing click — driver, executors, jobs, stages, tasks, and shuffles.

Read More
July 24, 2026
Databricks Troubleshooting

How to Update Nested (Struct) Columns in PySpark & Spark

withColumn("address.city", …) doesn't update the nested field — it quietly creates a weird new one. Here's how to actually update, add, and drop fields inside a struct.

Read More
July 24, 2026
1 2 3 4 Next »

Recent Posts

  • Liquid Clustering vs Z-ORDER vs Partitioning: Which to Choose
  • How Delta Lake Does ACID Without a Database
  • Spark Driver OOM: Why collect() and Broadcast Joins Crash Your Job
  • Delta Table Internals Explained: Files, Transaction Log, ACID & Time Travel
  • Medallion Architecture Explained Simply: Bronze, Silver & Gold

Archives

  • July 2026

Categories

  • Apache Spark
  • Data Engineering Basics
  • Databricks
  • Databricks Troubleshooting
  • Delta Lake
ByDataFix

All about data engineering

© 2026 ByDataFix. All rights reserved.

Subscribe to ByDataFix

Practical data engineering — PySpark, Azure, Snowflake and more. New posts straight to your inbox. No spam, unsubscribe anytime.