Zum Hauptinhalt springen
Podcasts & Audio BriefingsJohn Ternus is on a tear at Apple! [Cult of Mac podcast 35](03.10.2026 um 16:30 Uhr)
••
AI & KI Nachrichten"Muse Gadgets" turns AI hardware into an open-source DIY project(03.10.2026 um 16:28 Uhr)
•
AI & KI NachrichtenAn OpenAI safety employee has quit and is sounding the alarm(03.10.2026 um 16:31 Uhr)
••
AI & KI NachrichtenAll the AI agents that can live in your text messages(03.10.2026 um 16:00 Uhr)
•
AI & KI NachrichtenSpotify billionaire’s body scan startup has come to America(03.10.2026 um 16:00 Uhr)
•
AI & KI NachrichtenVessev built an electric ferry that almost flies(03.10.2026 um 16:42 Uhr)
•
AI & KI NachrichtenMeasuring the Creativity Potential of LLM Agents(03.10.2026 um 16:00 Uhr)
••
Podcasts & Audio BriefingsJohn Ternus is on a tear at Apple! [Cult of Mac podcast 35](03.10.2026 um 16:30 Uhr)
••
AI & KI Nachrichten"Muse Gadgets" turns AI hardware into an open-source DIY project(03.10.2026 um 16:28 Uhr)
•
AI & KI NachrichtenAn OpenAI safety employee has quit and is sounding the alarm(03.10.2026 um 16:31 Uhr)
••
AI & KI NachrichtenAll the AI agents that can live in your text messages(03.10.2026 um 16:00 Uhr)
•
AI & KI NachrichtenSpotify billionaire’s body scan startup has come to America(03.10.2026 um 16:00 Uhr)
•
AI & KI NachrichtenVessev built an electric ferry that almost flies(03.10.2026 um 16:42 Uhr)
•
AI & KI NachrichtenMeasuring the Creativity Potential of LLM Agents(03.10.2026 um 16:00 Uhr)
••
Intelligence View
⚡ tsecurity.de Intelligence

Spark Optimization

Scenario -1: joining tables, when one table is big and another is small Normal (Shuffle) join: All data move across the network and performance will be…

Beitrag
0
Seite
0
↗ Quelle (dev.to)
Social ReaktionenReagiere als Erste:r — dein Feedback zählt!




Scenario -1: joining tables, when one table is big and another is small




Normal (Shuffle) join: All data move across the network and performance will be degrade.




broadcast join



To avoid shuffle (slow network) developer can use broadcast join when on table table is small enough to fit in memory. Instead of shuffle the big table, broadcast join copy the small table to all executors/ clusters and big table reside as it it.

Finally join happen locally.




  • spark can auto broadcast if table size is below.



spark.sql.autoBroadcastJoinThreshold = 10 MB(Default)






  • Forcefully broadcast



from pyspark.sql.functions import broadcast
df = big_df.join(broadcast(small_df), "id)







Scenario -2: joining tables, when both tables are big.




When both tables are big, dev cannot use broadcast join, so Spark will do a shuffle join.

Spark must shuffle Table A by join key and shuffle Table B by join key and bring same keys to same executor.

This is slow + heavy network + also can cause of data skew.




1. repartition



Partition Both Tables on Join Key



If both tables are already partitioned by the same join column, then Spark shuffles once and keeps matching keys together




Helps balanced distribution.

Best when data is reused many times: write tables partitioned by join key.







df1 = df1.repartition("customer_id")
df2 = df2.repartition("customer_id")






2. Salting



Fix Data Skew



Some keys are very frequent (like country = "India")

One partition becomes huge → one task slow → job slow.

Add random value to skewed key:




df1 = df1.withColumn("salt", rand()*10)
df2 = df2.withColumn("salt", rand()*10)






Join using:




(key, salt)







Big key split into many small pieces

Work distributed across executors


Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Spark Optimization

Thematisch verwandte Begriffe: Spark, Optimization · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

💬 Kommentare werden geladen…
Zum Aktualisieren ziehen
Nächster Beitrag