1:1 mentoring with Big Tech AI engineers
System Design

Service and Container

Make a pod honest: readiness that means ready now, one trace id per request, a drain that waits for long runs with a derived grace period, and an image you can promote by digest.

Last updated

Production12 min readFirst readDeploy Your First Agent

After this section you can

  • Derive a grace period from a run's p99 and the preStop time, and say what it kills
  • Write liveness and readiness probes that fail for the right reasons
  • Build a digest-pinned, non-root image from a hashed lockfile
26

Service and Container

Readiness that means ready now, one trace id per request, a drain that waits for long runs, and one digest-pinned image.

Key idea

A service stays up while you replace it. Readiness tells the truth, the drain waits for the runs it holds, and the image ships pinned by digest. Agent runs can take tens of seconds, so the grace period is derived from the p99.

One request, one trace id, from the edge to the last log line
Edge mints the trace id traceparent Pod readiness gate, then the run Model call 529 when overloaded 1 · request Three log lines, one id each same id

Blue is the request, amber the decisions and the model call, purple the state, teal the trace id. Pans sideways on a phone.

Related

More in System Design

Get full access to all 74+ sections with code examples, diagrams, and interactive animations.

Unlock Premium