Skip to content
ByDataFix
  • Home
  • Contact
  • Privacy Policy
  • Terms & Conditions
Subscribe
Subscribe
Apache Spark

repartition vs coalesce in Spark: Which to Use and When

Both change the number of partitions — but one does a full shuffle and one doesn't, and that single difference decides your performance. Plus the coalesce(1) trap everyone falls into.

Read More
July 24, 2026
Apache Spark

Spark Architecture Explained Simply: Driver, Executors, Jobs, Stages & Tasks

The words all sound the same and that's why it's confusing. One kitchen analogy makes the whole thing click — driver, executors, jobs, stages, tasks, and shuffles.

Read More
July 24, 2026
Databricks Troubleshooting

How to Update Nested (Struct) Columns in PySpark & Spark

withColumn("address.city", …) doesn't update the nested field — it quietly creates a weird new one. Here's how to actually update, add, and drop fields inside a struct.

Read More
July 24, 2026
Databricks Troubleshooting

Fix Spark Data Skew in Joins (One Task Runs Forever)

Your job is "99% done" for an hour because one task got all the data. That's skew. Here's how to fix it with AQE, broadcast joins, skew hints, and salting.

Read More
July 24, 2026
Databricks Troubleshooting

Fix Corrupted Parquet Files in Spark (“Could Not Read Footer”)

One bad file shouldn't take down the whole read. Here's why Parquet files corrupt, how to skip past them to get your data now, and how to find and fix the real culprit.

Read More
July 22, 2026
Databricks Troubleshooting

Fix Delta ConcurrentAppendException (Concurrent Write Conflicts)

Two jobs wrote the same Delta table and one blew up. Here's what optimistic concurrency is really doing, and how to fix conflicts with partition filters, retries, and row-level concurrency.

Read More
July 22, 2026
Databricks Troubleshooting

Fix Over-Partitioned Delta Tables (Slow Queries) in Databricks

Partitioning by the wrong column shatters your table into millions of tiny files and grinds queries to a halt. Here's how to right-size it — and why liquid clustering is usually the better answer.

Read More
July 22, 2026
Databricks Troubleshooting

Why Empty Strings Become Null in a Partitioned Column (Spark)

No error, just wrong data: partition by a column with empty strings and Spark reads them back as null. Here's exactly why — and how to keep the values distinct.

Read More
July 22, 2026
1 2 3 Next »

Recent Posts

  • repartition vs coalesce in Spark: Which to Use and When
  • Spark Architecture Explained Simply: Driver, Executors, Jobs, Stages & Tasks
  • How to Update Nested (Struct) Columns in PySpark & Spark
  • Fix Spark Data Skew in Joins (One Task Runs Forever)
  • Fix Corrupted Parquet Files in Spark (“Could Not Read Footer”)

Archives

  • July 2026

Categories

  • Apache Spark
  • Data Engineering Basics
  • Databricks Troubleshooting
ByDataFix

All about data engineering

© 2026 ByDataFix. All rights reserved.

Subscribe to ByDataFix

Practical data engineering — PySpark, Azure, Snowflake and more. New posts straight to your inbox. No spam, unsubscribe anytime.