System Design

12) Compute & runtime — detailed guide

Concurrency models, pools, serverless vs containers, Kubernetes sizing, and safe runtime controls — explained with a Facebook-like Home Feed example.

A) Threads vs async I/O; pools; N+1 pitfalls

Threads vs async (how to choose)

  • Threads (blocking model): simplest mental model; good when blocking calls dominate and concurrency is moderate. Downsides: context switching + memory per thread.
  • Async/event loop: handles massive I/O concurrency (many sockets) if handlers remain non-blocking. Downsides: one accidental blocking call can stall everything (p99 spikes).
  • Hybrid is common: event loop for I/O + worker pool for CPU-heavy work (compression, crypto, image resizing).
  • Modern runtimes: Go goroutines, Java virtual threads, and async frameworks reduce pain, but you still must bound concurrency and pools.
Thread pool sizing — template (rule of thumb)
Template
threads ≈ cores × (1 + wait_time / compute_time)

Examples:
• Mostly CPU: wait/compute ~ 0 → threads ≈ cores
• 80% waiting, 20% compute: wait/compute = 4 → threads ≈ cores × 5

Always:
• Keep queues bounded
• Fail fast (429/503) once buffers are full
• Separate pools per dependency (DB vs external HTTP) to avoid head-of-line blocking

Connection pools (DB / HTTP / gRPC)

  • Goal: protect backends by bounding in-flight calls (your system survives spikes).
  • DB pool math: pool_per_pod × replicas ≤ DB max connections (leave headroom for admin, migrations, replicas).
  • HTTP: keep-alives on; prefer HTTP/2 multiplexing; cap per-origin concurrency.
  • Timeout layering: connect timeout, read timeout, overall deadline — all must fit within endpoint SLO.
Timeout & retry budget — template
Template
If your API SLO is 800ms: • Gateway budget: 50ms • Service compute: 200ms • DB budget: 250ms • Downstream budget: 250ms • Total: 750ms (leave margin) Retries: • Only for idempotent ops (or with idempotency key) • Total retry budget ≤ remaining time (don’t exceed SLO) • Use exponential backoff + full jitter

N+1 pitfalls (DB and microservices)

  • DB N+1: 1 query for the list + 1 query per row (feeds are a classic).
  • Service N+1: 1 feed call + N profile calls, or GraphQL resolvers doing per-item lookups.
  • Fixes: batch APIs, joins/IN queries, materialized read models (CQRS), GraphQL DataLoader, caching with bounded TTL.
Facebook-like Home Feed — runtime choices
Example
For GET /v1/feed: • Feed read is I/O-heavy: async/non-blocking clients (HTTP/2) + bounded concurrency. • CPU tasks (ranking / feature computation) run in a bounded worker pool. • DB pool per pod is small; total across pods fits DB max connections. • N+1 avoided by: - batched user/profile fetch - precomputed “fanout read model” or “timeline table” - GraphQL DataLoader if using GraphQL
Pitfalls
  • Blocking inside an event loop (p99 spikes and timeouts)
  • Unbounded queues that “hide” overload until you OOM
  • Fleet-wide DB pool explosion after scaling up or a deploy
  • Long timeouts that trigger retry storms during partial outages
  • N+1 hidden behind caches until the cache misses during a cold start

B) Containers vs serverless (cold starts, concurrency, cost)

Containers
  • Warm processes; predictable p99
  • Long-lived connections (streams, WebSockets)
  • Great for consumers, schedulers, stateful-ish services
  • You manage patching, scaling policy, bin packing
Serverless
  • Scale-to-zero; burst handling; per-request billing
  • Great for webhooks, cron-like tasks, glue code
  • Cold starts; max runtime; local disk is ephemeral
  • Concurrency caps and vendor limits matter
Serverless cold-start mitigation — template
Template
Mitigate cold starts: • Provisioned/warm concurrency for critical paths • Keep package/image small (trim deps, avoid huge layers) • Lazy-init heavy objects (ML models, large SDKs) • Reuse clients (HTTP pools, DB proxies) across invocations • Raise per-instance concurrency when safe (watch memory/CPU) Guardrails: • Set strict timeouts and concurrency limits • Use queues (SQS/Kafka) to smooth spikes into downstreams
Facebook-like Home Feed — where serverless fits
Example
• POST /v1/posts: image processing or fanout jobs can be serverless (async). • GET /v1/feed: keep in containers (latency + warm caches). • Use queues between the API and async workers to prevent burst overload.
Pitfalls
  • Assuming serverless is ‘infinite scale’ (account + regional caps exist)
  • Cold starts during sudden traffic spikes on user-facing endpoints
  • Using serverless for long-lived connections (WebSockets) without plan
  • Relying on ephemeral disk for durability

C) Kubernetes scheduling & bin packing (requests/limits, HPA, HA)

Requests vs limits (why it matters)

  • CPU: limits can cause throttling (CFS) → latency spikes. Prefer realistic requests; use limits carefully.
  • Memory: exceeding limit ⇒ OOMKill. Keep headroom; watch heap + off-heap + native.
  • QoS: Guaranteed (req=limit) vs Burstable (req<limit). Burstable is often best with good requests.
K8s sizing — template
Template
How I set requests: • CPU request ≈ p95 CPU usage (per pod) • Memory request ≈ p95 memory usage (per pod) + safety buffer • Avoid CPU limits on p99-sensitive services; if using limits, ensure you test throttling impact High availability: • Spread across zones (topologySpreadConstraints) • Pod disruption budgets for safe upgrades • Anti-affinity for critical services when needed

Autoscaling signals

  • APIs: scale on RPS/pod, in-flight, p95 latency, queue depth (custom metrics) — not CPU only.
  • Consumers: scale on lag + processing time; cap per-partition concurrency.
  • Coordination: scale safely with DB pools and downstream quotas.
Facebook-like Home Feed — K8s approach
Example
• Feed Service: HPA uses RPS/pod + p95 latency; stabilization prevents flapping. • Zone spread: keep replicas across AZs to survive an AZ failure. • Memory headroom avoids OOMKills during traffic spikes or cache warmups. • DB pool size is adjusted as replicas change so fleet_total ≤ DB max connections.
Pitfalls
  • Tight CPU limits throttling latency-critical pods
  • All replicas landing in one AZ (no topology spread)
  • HPA on CPU for I/O-bound services (wrong signal)
  • Scaling the API without scaling downstream capacity/pools

D) Service discovery, configuration & feature flags

Service discovery

  • Default: DNS-based discovery (Kubernetes Services) with stable virtual names.
  • Service mesh (optional): mTLS, retries, outlier ejection, traffic splitting, observability.
  • Client hygiene: deadlines/timeouts, jittered retries, and circuit breakers.

Configuration

  • Hierarchical config: defaults → environment → region/tenant → runtime flags.
  • Schema validation for config changes; hot reload only for safe settings.
  • Audit who changed what; canary config rollout before full blast.

Feature flags

  • Kill switches for risky features; percentage rollouts; per-tenant allowlists.
  • “Flag debt” policy: owner + expiry date; remove flags after rollout completes.
Facebook-like Home Feed — safe operations
Example
• Kill switch: disable “re-ranking v2” if p99 worsens, while keeping GET /v1/feed running. • Canary: route 1% of traffic to a new image pipeline, auto-rollback on error/latency deltas. • Config: DB pool sizing changes go through staged rollout + monitoring.
Pitfalls
  • Config drift and no audit trail
  • Leaking secrets via ConfigMaps or logs
  • Permanent feature flags increasing complexity over time
  • Blind retries in the mesh without budgets (cascading failures)

Handy rules & snippets (quote these in interviews)

Rules of thumb — compute & runtime
Template
Concurrency:
• I/O-bound → async multiplexing + bounded in-flight
• CPU-bound → bounded worker pool sized to cores and wait/compute ratio
• Never allow unbounded queues

Pools:
• DB pool sanity: pool_per_pod × replicas ≤ DB_max_conn
• HTTP: reuse connections, prefer HTTP/2 multiplexing, cap per-origin concurrency

Timeouts:
• Timeouts fit inside the endpoint SLO (layered budgets)
• Retries only within remaining budget, with full jitter

K8s:
• Requests ≈ p95 usage; keep memory headroom
• Avoid CPU limits on p99-sensitive services (throttling risk)
• HPA on meaningful signals: RPS/pod, p95, queue depth; stabilization windows

Flags:
• Every risky feature behind a kill switch
• Remove flags after rollout (flag debt policy)
Download on the App Store
Get it on Google Play
Thank You!

Your comments and complements are always the inspiration that keeps us motivated.



Address
  • Sterling, Virginia, 20166, US

  • info@techventurellc.com

  • 641 451 3463

  • 641 451 3463


Follow Us