Sathus AI 2.0 is now generally available — evaluation harnesses and guardrails included. Explore
Compute container allocation, off-heap buffers, and AQE partition targets based on Sathus distributed systems sizing formulas.
# Generated by Sathus Distributed Systems Calculator
spark.executor.instances: 16
spark.executor.cores: 4
spark.executor.memory: 10g
spark.executor.memoryOverhead: 3584m
spark.sql.shuffle.partitions: 2000
spark.sql.adaptive.enabled: true
spark.sql.adaptive.skewJoin.enabled: true
spark.sql.adaptive.coalescePartitions.enabled: true
spark.memory.fraction: 0.60
spark.memory.storageFraction: 0.50Assigning more than 5 cores to an executor leads to severe JVM stop-the-world garbage collection pauses. Keeping executor cores at 4 maximizes multi-threading while keeping GC pauses under 200ms.
PySpark delegates DataFrame transformations to Python worker processes (PyArrow C++ memory, Pandas UDFs). These run outside the JVM. Without 20-25% memoryOverhead, YARN or K8s cgroups will send SIGKILL (Exit code 137).
Spark tasks achieve optimal throughput when processing 100MB to 200MB per partition. Smaller partitions create metadata scheduling overhead; larger partitions (>2GB) crash the Spark internal ByteBuffer limit.