Service and Container
Make a pod honest: readiness that means ready now, one trace id per request, a drain that waits for long runs with a derived grace period, and an image you can promote by digest.
Last updated
After this section you can
- Derive a grace period from a run's p99 and the preStop time, and say what it kills
- Write liveness and readiness probes that fail for the right reasons
- Build a digest-pinned, non-root image from a hashed lockfile
Service and Container
Readiness that means ready now, one trace id per request, a drain that waits for long runs, and one digest-pinned image.
A service stays up while you replace it. Readiness tells the truth, the drain waits for the runs it holds, and the image ships pinned by digest. Agent runs can take tens of seconds, so the grace period is derived from the p99.
Blue is the request, amber the decisions and the model call, purple the state, teal the trace id. Pans sideways on a phone.