2026-07-29
Separate Kubernetes probe shutdown grace from Pod deletion
Size Spring Boot shutdown against distinct Pod-deletion and probe-failure grace periods, then validate and drill each path separately.
One Pod field controls ordinary deletion. An optional field in a startup or liveness probe can select a different clock only when that probe causes a container termination. Treating those clocks as interchangeable is how an otherwise reasonable Spring shutdown budget becomes either a forced stop or an unnecessarily slow recovery.
This is source-reviewed guidance, reviewed on 2026-07-29 against Spring Boot 4.1 and current Kubernetes documentation. No Kubernetes cluster or Spring Boot application was run for this publication. No server-side dry run, ordinary deletion drill, or liveness-failure drill was executed. All YAML, commands, values, expected observations, and acceptance tables are operator templates. The 20-second Spring phase timeout, 30-second Pod grace, and 25-second liveness grace below are illustrative only.
Name the two termination paths
The trigger decides both the identity outcome and the applicable Kubernetes grace period. The Kubernetes probe guide distinguishes failed startup and liveness probes, which terminate or restart a container, from a failed readiness probe, which makes the container unready and continues probing it. The Pod lifecycle documentation describes the separate Pod termination path.
| Trigger | Kubernetes reaction | Grace source | Identity outcome |
|---|---|---|---|
| API requested Pod deletion or rollout replacement | The Pod terminates | Pod-level terminationGracePeriodSeconds, or the documented 30-second default |
The old Pod is deleted; a controller can make a replacement with a new Pod UID |
| Startup probe failure | The failed container terminates and follows its restart policy | startupProbe.terminationGracePeriodSeconds when set, otherwise the Pod-level value |
A container can restart in the same Pod |
| Liveness probe failure | The failed container terminates; it restarts only when the Pod restart policy permits it | livenessProbe.terminationGracePeriodSeconds when set, otherwise the Pod-level value |
The Pod UID stays; after a permitted restart, the container ID and restart count change |
| Readiness probe failure | The container is marked unready and continues running | No termination grace applies | The container keeps running while ordinary Service traffic should stop routing to it |
ordinary deletion:
API deletion -> Pod grace starts -> optional preStop -> TERM -> Spring shutdown -> forced stop after the nominal budget
liveness failure:
probe failures reach threshold -> probe grace starts -> optional preStop -> TERM -> Spring shutdown -> forced stop after the nominal budget -> permitted container restartProbe detection happens before either termination clock starts. periodSeconds, timeoutSeconds, and failureThreshold determine when the failure becomes actionable; they are not part of the Pod-level or probe-level grace period. A readiness failure is not a short liveness failure.
Kubernetes can grant a small one-off two-second extension when preStop is still running as the grace period expires. Treat that as an exceptional cleanup allowance, not capacity in the shutdown budget. Size preStop, Spring shutdown, and margin to finish within the configured grace period itself.
During Pod termination, a terminating EndpointSlice endpoint is represented as not ready. That routing state is useful evidence that ordinary traffic should drain, but it is not evidence that the JVM has stopped. The Kubernetes documentation does not establish the precise behavior of every external load balancer, so measure that separately in the environment that serves the traffic.
Bound Spring shutdown before choosing numbers
Spring Boot 4.1 enables graceful shutdown by default for embedded Jetty, Reactor Netty, and Tomcat, in servlet and reactive applications. Context closure gives existing requests a chance to complete while those servers stop accepting new requests at the network layer. See the Spring Boot 4.1 graceful shutdown reference.
The same reference describes spring.lifecycle.timeout-per-shutdown-phase as a timeout for each shutdown phase. A per-phase Spring timeout is not a whole-process shutdown guarantee. Several lifecycle phases can consume time, and shutdown can also include bean destruction, executor closure, connection cleanup, and application-owned stop work.
The following is an unexecuted operator template, not a recommended setting:
spring:
lifecycle:
timeout-per-shutdown-phase: "20s"
management:
endpoint:
health:
probes:
add-additional-paths: trueThe second property exposes the liveness and readiness groups on /livez and /readyz at the main server port. That avoids a management listener reporting healthy when the application listener cannot accept connections. Spring Boot's Kubernetes probe reference documents those paths and advises against making liveness depend on shared external systems. Readiness dependencies are a separate failure-policy choice.
Use measured terms instead of translating the 20-second property directly into a Kubernetes value:
P = measured maximum preStop duration for this path
S = measured upper bound from TERM to application process exit
M = explicit runtime, observation, and scheduling margin
Gpod = Pod-level grace for ordinary deletion
Gprobe = startup or liveness probe-level grace for probe-triggered termination
Gpod > P + S + M
Gprobe > P + S + MThe applicable grace countdown begins before preStop, so the hook consumes the same budget rather than adding time before it. Derive S from the lifecycle-phase inventory and application-owned stop behavior, then measure the whole TERM-to-exit interval repeatedly. The unexplained remainder is not a margin. If a deliberately shorter Gprobe cannot contain that work, it is an accepted forced-termination risk intended to accelerate recovery, not graceful shutdown.
| Setting | Illustrative value | Meaning |
|---|---|---|
| Spring timeout per shutdown phase | 20s |
Bounds each applicable phase, not total process exit |
| Pod-level grace | 30s |
Nominal shutdown budget for ordinary Pod deletion; do not budget against the exceptional preStop extension |
| Liveness probe-level grace | 25s |
Nominal shutdown budget for a liveness-triggered termination when the field is present; do not budget against the exceptional preStop extension |
20 < 25 < 30 does not make either path fit. The configuration becomes defensible only when the measured P + S + M is below both configured Kubernetes budgets, without relying on the exceptional extension.
Put each field where the API permits it
This complete Deployment is an unexecuted operator template. Its registry name and digest are deliberately invalid placeholders, not a runnable canary artifact.
apiVersion: apps/v1
kind: Deployment
metadata:
name: spring-probe-shutdown-canary
spec:
replicas: 1
selector:
matchLabels:
app: spring-probe-shutdown-canary
template:
metadata:
labels:
app: spring-probe-shutdown-canary
spec:
terminationGracePeriodSeconds: 30
containers:
- name: app
image: "registry.example.invalid/spring-probe-shutdown-canary@sha256:<digest>"
ports:
- name: http
containerPort: 8080
startupProbe:
httpGet:
path: /livez
port: http
periodSeconds: 2
failureThreshold: 30
livenessProbe:
httpGet:
path: /livez
port: http
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 3
terminationGracePeriodSeconds: 25
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 2Under spec.template.spec, the Pod-level field is terminationGracePeriodSeconds. The probe-level field belongs directly inside startupProbe or livenessProbe, alongside failureThreshold, never inside httpGet. This template's startup probe inherits the Pod-level grace because it has no override. Do not put the field in readinessProbe; readiness does not terminate the container and Kubernetes rejects that placement.
The general liveness row above is conditional on restartPolicy. This Deployment's regular app container uses the default restartPolicy: Always, so this drill expects a same-Pod restart, a changed container ID, and an increased restart count.
The probe-level termination grace documentation says the feature was available from Kubernetes 1.25 and is stable from 1.28. It also documents a minimum probe-level value of one second. Generated API prose can retain older beta or feature-gate wording, so use the concept guide for feature state and the Pod Probe API plus the target API server for actual support. Do not use zero as an immediate probe-shutdown template.
Ask the target API server
The commands below are unexecuted operator templates. kubectl explain makes the served schema paths reviewable. A server-side dry run asks the real API server to decode, default, validate, and run applicable admission without persisting the object. A local YAML parser or client-side dry run is useful for syntax, but it cannot substitute for the target server.
kubectl version
kubectl explain pod.spec.terminationGracePeriodSeconds
kubectl explain pod.spec.containers.startupProbe.terminationGracePeriodSeconds
kubectl explain pod.spec.containers.livenessProbe.terminationGracePeriodSeconds
kubectl explain pod.spec.containers.readinessProbe
kubectl apply --server-side --dry-run=server \
--filename spring-probe-shutdown-canary.yamlRecord the actual cluster version, client version, manifest digest, command, and API response when an operator runs this. Do not invent acceptance output. This narrow review catches a common field-placement mistake without applying an invalid object to a shared environment:
rg -n -C 3 'terminationGracePeriodSeconds' \
spring-probe-shutdown-canary.yamlThe review expectation is one Pod-level field and one liveness probe-level field, with none under readiness or httpGet.
Drill ordinary Pod deletion
Run this only in a disposable namespace with a single-replica canary Service and Deployment. Prepare an immutable image, the exact Spring Boot version, a production-like non-sensitive request that lasts long enough to observe draining, and logs that survive Pod-object deletion. Record the actual preStop content and its measured upper bound. Do not use --force or --grace-period=0.
Terminal one captures identity and watches lifecycle state. This is an unexecuted operator template:
NS='<canary-namespace>'
APP_LABEL='app=spring-probe-shutdown-canary'
POD="$(kubectl --namespace "$NS" get pods \
--selector "$APP_LABEL" \
--output jsonpath='{.items[0].metadata.name}')"
OLD_UID="$(kubectl --namespace "$NS" get pod "$POD" \
--output jsonpath='{.metadata.uid}')"
printf 'old_uid=%s\n' "$OLD_UID"
kubectl --namespace "$NS" get pod "$POD" \
--output custom-columns='NAME:.metadata.name,CONTAINER_ID:.status.containerStatuses[?(@.name=="app")].containerID,RESTARTS:.status.containerStatuses[?(@.name=="app")].restartCount'
kubectl --namespace "$NS" get pod "$POD" --watch --output wideTerminal two starts the declared long request. Keep it separate from the new-traffic observer so an already accepted request is not mistaken for a new request. This is an operator template:
curl --fail-with-body --no-buffer \
--max-time '<request-timeout-seconds>' \
'https://<canary-service-address>/<bounded-long-request-path>'Terminal three watches the canary Service's EndpointSlice conditions and continuously samples new ordinary traffic through that Service. Stop the traffic loop manually after the observation window. These are operator templates:
NS='<canary-namespace>'
SERVICE='spring-probe-shutdown-canary'
kubectl --namespace "$NS" get endpointslices \
--selector "kubernetes.io/service-name=$SERVICE" \
--watch \
--output custom-columns='SLICE:.metadata.name,POD:.endpoints[*].targetRef.name,UID:.endpoints[*].targetRef.uid,READY:.endpoints[*].conditions.ready,SERVING:.endpoints[*].conditions.serving,TERMINATING:.endpoints[*].conditions.terminating'while :; do
printf '%s ' "$(date --utc +'%Y-%m-%dT%H:%M:%S.%3NZ')"
curl --silent --show-error --fail-with-body \
--write-out '\nhttp_status=%{http_code} total_seconds=%{time_total}\n' \
'https://<canary-service-address>/<ordinary-request-path-that-returns-serving-pod-uid>' || \
printf 'curl_exit=%s\n' "$?"
sleep '<sampling-interval-seconds>'
doneThe ordinary response must include the serving Pod UID, obtained by the canary application from the downward API or an equivalent immutable startup value. Status and timing alone cannot show whether the old endpoint received a request.
Terminal four deletes normally, observes rollout completion, and waits only a bounded time for a Ready replacement whose UID differs from OLD_UID. It discovers its own variables because shell state does not cross terminal sessions. This is an operator template:
NS='<canary-namespace>'
APP_LABEL='app=spring-probe-shutdown-canary'
POD="$(kubectl --namespace "$NS" get pods \
--selector "$APP_LABEL" \
--output jsonpath='{.items[0].metadata.name}')"
OLD_UID="$(kubectl --namespace "$NS" get pod "$POD" \
--output jsonpath='{.metadata.uid}')"
date --utc +'%Y-%m-%dT%H:%M:%S.%3NZ'
kubectl --namespace "$NS" delete pod "$POD" --wait=false
kubectl --namespace "$NS" rollout status \
deployment/spring-probe-shutdown-canary \
--timeout='<rollout-timeout>'
replacement_wait_seconds='<replacement-wait-seconds>'
deadline=$((SECONDS + replacement_wait_seconds))
ready_replacement=''
while (( SECONDS < deadline )); do
replacement="$(kubectl --namespace "$NS" get pods \
--selector "$APP_LABEL" \
--field-selector=status.phase=Running \
--output custom-columns='NAME:.metadata.name,UID:.metadata.uid,READY:.status.conditions[?(@.type=="Ready")].status' \
--no-headers)"
ready_replacement="$(awk -v old_uid="$OLD_UID" \
'$2 != old_uid && $3 == "True" { print $1, $2; exit }' \
<<< "$replacement")"
if [[ -n "$ready_replacement" ]]; then
printf 'ready_replacement=%s\n' "$ready_replacement"
break
fi
sleep 1
done
if [[ -z "$ready_replacement" ]]; then
printf 'no Ready replacement with a UID different from OLD_UID within %ss\n' \
"$replacement_wait_seconds" >&2
exit 1
fi
kubectl --namespace "$NS" get events \
--field-selector "involvedObject.name=$POD" \
--sort-by='.lastTimestamp'The acceptance condition is evidence that the deleted Pod used the Pod-level grace, not the liveness override; the EndpointSlice observer records the old endpoint's readiness, serving, and terminating state; timestamped Service traffic shows when new ordinary requests stop reaching the old endpoint; and the already accepted bounded request has either met its stated contract or has a documented failure. Readiness alone cannot establish request completion. The bounded loop records a Ready replacement with a Pod UID different from OLD_UID rather than assuming rollout completion proves replacement identity.
Collect timestamps for deletion to exit and TERM to exit separately where the environment exposes both. Logs should show TERM, Spring context shutdown, bounded phase behavior, and process exit before the Pod-level deadline. There must be enough evidence to distinguish a graceful exit from a forced stop. If the platform does not retain it, the result is inconclusive. The replacement should have a new Pod UID and become ready.
Drill liveness failure separately
Use a fresh canary and do not delete its Pod. The drill needs an explicit, pre-reviewed canary-only injector that makes /livez fail while process 1 remains alive to receive TERM. It must be absent or unreachable in production. kill -STOP 1, direct SIGKILL, Pod deletion, and node failure exercise different paths and cannot establish the liveness-to-grace behavior.
Capture the baseline first. This is an unexecuted operator template:
NS='<canary-namespace>'
POD='<canary-pod>'
CONTAINER='app'
kubectl --namespace "$NS" get pod "$POD" \
--output jsonpath='uid={.metadata.uid}{"\n"}containerID={.status.containerStatuses[?(@.name=="app")].containerID}{"\n"}restarts={.status.containerStatuses[?(@.name=="app")].restartCount}{"\n"}'
kubectl --namespace "$NS" get pod "$POD" --watchInvoke only the real canary mechanism. There is no universal liveness-failure injector, so this unresolved placeholder is intentional:
kubectl --namespace "$NS" exec "$POD" --container "$CONTAINER" -- \
<canary-only-command-that-makes-livez-fail>After the intended restart, collect the termination record and previous-container logs. This too is an operator template:
kubectl --namespace "$NS" describe pod "$POD"
kubectl --namespace "$NS" logs "$POD" \
--container "$CONTAINER" \
--previous \
--timestamps
kubectl --namespace "$NS" get pod "$POD" \
--output jsonpath='uid={.metadata.uid}{"\n"}containerID={.status.containerStatuses[?(@.name=="app")].containerID}{"\n"}restarts={.status.containerStatuses[?(@.name=="app")].restartCount}{"\n"}lastReason={.status.containerStatuses[?(@.name=="app")].lastState.terminated.reason}{"\n"}lastExitCode={.status.containerStatuses[?(@.name=="app")].lastState.terminated.exitCode}{"\n"}'The acceptance condition is that the configured failure threshold occurs before termination, with detection latency recorded separately from shutdown grace. The Pod UID remains the same, while the container ID changes and restart count rises by the intended number of restarts. Previous-container timestamps should show TERM and Spring shutdown within the liveness probe-level deadline. Missing or ambiguous termination records make the result inconclusive. The restarted container should recover /livez and /readyz, then the Pod becomes ready. Remove the injector or destroy the disposable canary after evidence collection.
| Evidence | Ordinary deletion | Liveness failure |
|---|---|---|
| Trigger time | Record | Record |
| Probe threshold reached | Not applicable | Record |
| Pod UID before and after | Changes when replacement appears | Same |
| Container ID before and after | Replacement has a new container | Changes in the same Pod |
| Grace selected | Pod-level | Liveness probe-level |
preStop duration |
Record | Record |
| TERM to Spring exit | Record | Record |
| Forced-stop evidence | Must be absent or inconclusive | Must be absent or inconclusive |
| Existing request outcome | Record | Record if included |
| Recovery outcome | Replacement ready | Restarted container ready |
Keep HTTP graceful shutdown in its lane
Graceful embedded-server shutdown can bound handling for accepted HTTP requests. No claim is made that graceful HTTP shutdown makes messages, outbox events, background jobs, executor work, database effects, or external side effects safe. Each non-HTTP path needs its own stop-intake behavior, bounded work, durable ownership, idempotency, and recovery drill.
A passing readiness check does not drain those paths. A longer probe-level grace can delay liveness recovery; a shorter one can force incomplete work. Choose any difference between Pod-level and probe-level grace from a measured service objective, not from the illustrative values in this template.