Skip to content

Wake order

Describes Hibernator chart 0.12.44

Some workloads must be up before others can start: a service mesh before the apps in it, a database before its clients. Wake tiers set that order, and hibernate tiers set the order of the scale-down. This page also says when a wake stops waiting for a workload, and when a wake is released and when it settles. Every value named here is in the table in configuration.md.

A wake tier orders the wake: a resource in tier N scales up only when every tier above it is ready. Tiers are released in descending priority. Everything without a priority stays in one unordered pass. The tier list is for charts that you do not own. A hibernator.io/wake-priority annotation on the Deployment or StatefulSet wins over the list (see annotations.md).

config:
wakeOrdering:
tiers:
- priority: 100 # >= 1 and unique
timeout: "20m" # optional, default: config.controller.wakeBarrierTimeout
targets: # optional: a target-less tier only sets a timeout
- namespace:
exact: "linkerd" # or: pattern: "mesh-*"
waitOnly: true # read readiness, never change replicas

There is no default tier list, on purpose: where a resource goes is your decision. A database joins a tier through the wakePriority of its entry (see databases.md). Put the database in the same tier as the service mesh, not above it. Members of one tier start in parallel. A database in its own tier above the mesh holds the mesh, and every app in the mesh, for the full instance start.

YAML anchors in wakeOrdering and hibernateOrdering follow the rule in databases.md: define the anchor in the same section.

A tier is ready when every pod of every member passes its readiness probe. A CronJob counts as ready at once. A settle delay of 2 seconds follows every tier.

A database is ready at RDS status available and at nothing earlier. The listener accepts connections minutes before, while RDS still restarts the engine.

A waitOnly target is read and never scaled. This lets a service mesh control plane go first when it is in a namespace that Hibernator does not manage.

A tier that holds a database and sets no timeout of its own uses 20 minutes, not config.controller.wakeBarrierTimeout. Its clock starts again on every forward step of stopping -> stopped -> starting -> available. A wake that meets a stop thus completes and does not fail open. Only a forward step counts. Otherwise a status that flaps would reset the clock on every pass, and the barrier could never fail open.

The wake’s own fail-open starts to count only when the last group is released. No barrier wait uses up the fail-open, so you do not have to increase a value for a database tier.

A tier that exceeds its timeout fails open. Hibernator then:

  • logs it,
  • sends a Wake Barrier Timeout notification that names the members that never became ready,
  • increments hibernator_wake_barrier_timeouts_total{tier} on the controller,
  • releases the remaining tiers.

Nodes are provisioned inside that window, so a cold cluster can need more than the 5-minute default. Hibernation has its own order, see Hibernate ordering.

Writing off a workload that will never start

Section titled “Writing off a workload that will never start”

A wake waits for every managed workload, so one workload that cannot start would hold the wake for ever. Hibernator judges each workload against its own declaration, in two phases:

  1. A flat 15 minutes for the cluster to produce a started container at all.
  2. The budget that its startupProbe declares: initialDelaySeconds + failureThreshold x periodSeconds, the longest among the containers that declare one, capped at 30 minutes.

A workload that declares no startupProbe gets config.controller.workloadSettleTimeout for the second phase. So does a workload whose spec Hibernator cannot read. The accepted range is 1m-30m.

Behind those deadlines, the whole transition has a fail-open. It is derived from the same declarations, not configured. It is the highest ceiling that any of your workloads can reach, plus two minutes, and never less than 15 minutes. There is no value to set. Because it comes from the same numbers, it cannot disagree with the per-workload deadlines that it bounds.

A wake can still correctly run longer than that figure. The fail-open waits while a written-off workload is inside its grace band. It also waits while the kubelet has to kill a workload that was written off on its declared startupProbe budget. It then fires at the later of two times. The first is the derived deadline. The second is the end of the latest live band or kill wait, plus a margin of 45 seconds. The outside figure is thus the derived fail-open plus 3m45s, or plus 5m45s for a wake that a controller restart resumed. Use the outside figure when you size a runbook.

The clock is per workload and starts when that workload was scaled, not when the transition started. A workload in the third wake group must not get a clock that ran down while the first two groups came up.

Hibernator writes off a workload before its deadline only in three cases:

  • An image reference that is not a valid reference.
  • An image that is not on the node, with pull policy Never.
  • A container that restarted twice before it was ready.

No clock can improve the first two. Init containers count for them too: a pod whose init container cannot pull its image never reaches its app containers.

A failed pull is not among them, on purpose. The kubelet backs off from 10 seconds, so ErrImagePull and ImagePullBackOff both appear within seconds of one failed pull, and then alternate. The pair cannot tell a short registry failure from a broken image. A write-off for it would end the whole wake while the pod was about to come up. A pull that succeeds inside the 15 minutes costs the workload nothing. A pull that still fails at the ceiling is a phase-1 expiry that names the registry.

Unschedulable is not one of them either, on purpose. On a cold cluster with node autoscaling, every pod is unschedulable for the first minutes while the nodes are provisioned. A write-off would report a failure for a wake that works as designed. These pods wait out phase 1 like all others. Phase 1 is the flat 15 minutes, and it exists so that this wait happens first.

A written-off workload stays in the totals as a failure: the pod count that the UI shows never drops by one. Only the progress bar’s own denominator drops it, so the bar keeps moving on the pods that can still arrive.

Hibernator logs each write-off once, with a full sentence that names what ran out or what went wrong:

  • did not finish starting within its declared 20 minutes
  • was still Unschedulable after 15 minutes
  • its init containers had not finished after 15 minutes
  • could not pull its image in 15 minutes
  • could not create its container (…)
  • restarted twice before it was ready
  • could not be read

The activity log, the completion dialog, the Teams card, the email and the audit entry all show this sentence word for word.

A workload that Hibernator later sees running is no longer written off. An observation outranks the earlier verdict, and the wake reports the workload as recovered, not as a casualty. The activity log replaces its line in place. The recovery has its own sentence, which starts with the reason it lifted: did not finish starting within its declared 20 minutes, then came up less than 1 minute after the restart. The recovery shows in these places:

  • a Recovered block on the Teams card and a panel in the email, which both name the first three workloads and count the rest,
  • workloads_recovered on the audit entry,
  • a counted clause on the Kubernetes event,
  • the hibernator_workload_recoveries_total counter.

None of it escalates. A wake that lost nothing and recovered two workloads stays plain and green, and goes where a clean wake of its kind goes. Only a degraded wake also reaches the high-priority destination of this cluster. troubleshooting.md has the shapes of the sentence and what each means.

When a wake is released and when it settles

Section titled “When a wake is released and when it settles”

A wake releases when your environment answers and settles when every workload has a verdict. On a large cluster those are minutes apart, and the difference is what a user actually waits for.

The derived fail-open is the settle deadline: the moment Hibernator stops waiting for the whole transition, whatever its workloads still do. Its clock starts when the last group barrier releases, because the per-workload ceilings that it comes from start there. A wake that waited 20 minutes at a database group thus still gives every workload below it its full budget. The wake as a whole can correctly run much longer than the fail-open figure.

The origin-URL probe has a flat 75-minute backstop of its own. It exists only to catch a transition that nobody cleared.

The fail-open also waits for the grace bands and the kill waits of written-off workloads (see troubleshooting.md). The outside figure for a wake is the derived fail-open plus 3m45s, or plus 5m45s for a wake that a controller restart resumed.

The origin-URL probe observes the release, so release needs config.urlReadinessCheck.enabled: true and an origin URL for the probe (see url-readiness.md). An install without one releases when it settles, and nothing about such a wake changes.

When you have release, it changes two things:

  • Resource discovery resumes at release, not at settle, so the controller is not blind during the settling time.
  • The overlay counts down a wake-duration estimate that Hibernator learns from release durations, not settle durations. It thus estimates the wait that your users see.

Release does not change these:

  • The wake stays open until it settles.
  • A completed wake still writes its one audit entry, one Kubernetes event and one notification once, at settle.
  • Scheduled hibernation stays held for the whole transition, settle included.

A hibernate tier orders the scale-down. Hibernator scales a tier to zero and confirms that its pods are gone before it touches the tier below. Tiers go down in descending priority, and everything without a priority goes down last. The numbers are independent of wakeOrdering: one number cannot say “scale down first and wake up first”. The tier list is for charts that you do not own. A hibernator.io/hibernate-priority annotation on the Deployment, StatefulSet or CronJob wins over the list (see annotations.md).

config:
hibernateOrdering:
tiers:
- priority: 100 # >= 1 and unique
timeout: "2m" # optional, default: config.controller.hibernateTierTimeout
targets: # optional: a target-less tier only sets a timeout
- namespace:
exact: "keda" # or: pattern: "keda-*"

There is no default tier list, on purpose: where a resource goes is your decision. The case for hibernate tiers is KEDA. Its operator watches the deployments that Hibernator scales. In the seconds before its own pod stops, its scale loop writes them back to minReplicaCount. On a test cluster, this brought more than a third of the deployments back to one replica after a hibernation had reported complete. Put the keda namespace in the first tier to remove that window. Hibernator itself learns no KEDA resource type, on purpose.

A tier is done when every pod of the resources it still owes the tiers below is gone: gone, not Terminating. Then the same 2-second settle delay as in the wake barrier follows. These resources are what the tier itself scaled down, plus anything that an earlier pass already set to zero whose pods have not gone. A walk that stops in the middle of a drain leaves exactly that behind. The walk of the next reconcile must continue from that boundary.

Hibernator does not wait for:

  • a resource whose scale-down failed, because its pod count could never reach zero,
  • a resource that an external wake session holds up, because its pods come back and do not go,
  • a CronJob, which has no pods.

Unlike the wake barrier, the whole order runs inside one scaling pass. A hibernate transition thus holds the reconcile loop for the length of its tier waits. Keep the per-tier timeout short, as the 2m of the KEDA tier above. Do not use the 5-minute default.

A tier that exceeds its timeout fails open. Hibernator logs it, records it on the hibernation audit entry and on the HibernationActivated event, and scales the remaining tiers anyway. The worst case is then the drift of an unordered hibernation, which the hourly reconcile still corrects.

Managed databases are not in the hibernate tiers, on purpose: the stop side has its own stopDelay and minimumStopDuration (see databases.md).