Skip to content

Annotations

Describes Hibernator chart 0.12.44

Annotations you put on your own workloads and namespaces to control what Hibernator does with them. All of them use the hibernator.io/ prefix.

This is the file you want if you need one workload left alone: put hibernator.io/exclude: "true" on it and Hibernator will never scale it.

A CronJob you suspended yourself needs no annotation to stay suspended: Hibernator records a suspension it finds as yours, and a wake leaves it in place. Lift it and the CronJob follows the schedule again. Hibernator cannot tell your suspension from its own, and takes it for its own, when it finds it while it holds the CronJob suspended for hibernation, after a hibernation or wake of its own was interrupted, or within an hour of its own last suspend or resume. It resumes such a suspension on its next pass if the CronJob should be running, and at the next wake otherwise. To keep the CronJob suspended, suspend it again: once Hibernator finds the suspension more than an hour after its own last resume, it records it as yours.

Hibernator also writes annotations of its own onto the resources it manages, to track wake state, pod protection and redirect state. Those are internal, they are not listed here, and setting them by hand will confuse the controller.

Exclude a resource or namespace from hibernator management.

  • Target: Namespace, Deployment, StatefulSet, CronJob
  • Values: "true"
  • Precedence: Wins over all other annotations (safety first)
metadata:
annotations:
hibernator.io/exclude: "true"

Include a resource in an excluded namespace. Allows opting-in individual resources when the namespace is excluded.

  • Target: Deployment, StatefulSet, CronJob
  • Values: "true"
  • Note: Only meaningful when the namespace is excluded
metadata:
annotations:
hibernator.io/include: "true"

Order the wake. A resource scales up only once every tier above it is ready.

  • Target: Deployment, StatefulSet, CronJob
  • Values: a positive integer. Tiers are released in descending order, so 100 goes before 10. Absent, 0 or unparseable keeps the resource in the single unordered pass that is the default behavior.
  • Precedence: wins over the wakeOrdering.tiers list in the ConfigMap, which is the fallback for resources that carry no annotation
metadata:
annotations:
hibernator.io/wake-priority: "100"

How the barrier behaves:

  • A tier is ready when every pod of every member passes its readiness probe – readiness, not Running. A CronJob has no pods and counts ready at once, so it takes part in the order without ever blocking. A resource whose baseline is 0 (user-stopped) is nothing to wait for and counts ready too.
  • A fixed 2-second settle delay follows every tier: a pod is ready shortly before the endpoint controller publishes its address.
  • Each tier has a timeout – its own timeout in the tier list, otherwise controller.wakeBarrierTimeout (default 5m), or 20m for a tier that holds a managed database. Node provisioning happens inside that window, so a cold cluster can need more than the default.
  • On timeout hibernator fails open: it logs the tier, sends a Wake Barrier Timeout notification naming the members that never became ready, increments hibernator_wake_barrier_timeouts_total, and releases the remaining tiers. An environment that refuses to wake because one dependency is slow is worse than a few pods that start early.
  • Hibernation has its own order, under hibernator.io/hibernate-priority below. The annotation carries wake- in its name because a single number read backwards cannot express “scale down first and wake up first”, so the two number spaces are independent.

A tier can also hold a managed database – an entry of the databases section, placed by its own wakePriority. It is ready at RDS status available and at nothing weaker, its clock restarts on every forward step of that sequence so a wake during a stop is not counted as a stall (a flapping status resets nothing), and a wake that finds the instance stopped starts it when the tier is released.

A tier can also hold a wait-only target – a namespace hibernator does not manage, named in the wakeOrdering.tiers list with waitOnly: true. Hibernator reads its readiness and never changes its replicas, which is how a service mesh control plane can go first. A wait-only target that is scaled to zero counts as ready.

Order the hibernation. A tier is scaled down and its pods confirmed gone before the tier below it is touched.

  • Target: Deployment, StatefulSet, CronJob
  • Values: a positive integer. Tiers are scaled down in descending order, so 100 goes before 10. Absent, 0 or unparseable leaves the resource in the unordered remainder, which goes down last.
  • Precedence: wins over the hibernateOrdering.tiers list in the ConfigMap, which is the fallback for resources that carry no annotation
  • Independent of wake-priority: the two number spaces do not have to agree, and neither implies the other
metadata:
annotations:
hibernator.io/hibernate-priority: "100"

How the walk behaves:

  • The whole order runs inside one scaling pass. A tier is done when every pod of the resources it just scaled down is gone – gone, not Terminating – plus the same 2-second settle delay the wake barrier uses. A CronJob has no pods, so suspending it never blocks, and neither do the satellites standing in for a redirected service: they carry the service’s selector but are not the workload’s pods, so the count excludes them.
  • A tier waits for what it owes the tiers below it: the resources this pass took down, plus any an earlier pass already zeroed whose pods have not gone away. The second kind is what a walk that stopped mid-walk leaves behind, and skipping it would take the tiers below down while the tier above them was still terminating. A resource whose scale-down failed is still not waited on – there would be no pod count that could ever reach zero – and neither is one an external wake session is holding up, whose pods are coming back rather than going.
  • Each tier has a timeout – its own timeout in the tier list, otherwise controller.hibernateTierTimeout (default 5m).
  • On timeout hibernator fails open: it logs the tier, records it on the hibernation audit entry and the HibernationActivated event, and scales the remaining tiers anyway. The worst case is then the drift of an unordered hibernation, which the hourly reconcile still corrects.
  • The transition overlay renders the wait as “group N of M”, with the same activity log the wake barrier uses, so a tier wait is distinguishable from a hung transition.
  • A tier wait runs inside the reconcile lock, so a Wake click (or any other hibernator-state change) arriving mid-walk ends the current wait and leaves the remaining tiers to the reconcile it triggered. The click is served within a poll interval instead of after the whole walk, and the abandoned tiers are not recorded as failing open.

The case this exists for is KEDA: its operator watches the very deployments hibernator scales, and its scale loop writes them back to minReplicaCount in the seconds before its own pod dies. Placing the keda namespace in the first tier removes the window. Hibernator itself deliberately learns no KEDA resource types. A chart hibernator does not own can never be annotated, which is why the tier for it comes from the ConfigMap.

Enable satellite redirect on a Service when the backing resource is hibernated.

  • Target: Service
  • Values: "true"
metadata:
annotations:
hibernator.io/redirect-when-hibernating: "true"

Custom port for satellite redirect.

  • Target: Service
  • Values: Port number as string (e.g., "8080")

Protocol for satellite redirect.

  • Target: Service
  • Values: "http", "https"
  • Certificate: with "https", the frontend serves the certificate from tls.secretName when tls.enabled is true. Otherwise it serves a certificate from the internal CA of Caddy, which browsers do not accept. See service-redirection.md.

Prevent hibernator from updating the stored OriginalReplicas baseline for a resource. When set, hibernator preserves the baseline even if live replica counts change.

  • Target: Deployment, StatefulSet
  • Values: "true" (any other value is ignored)
  • Use case: External tools that temporarily scale a workload, e.g. a test harness that pins KEDA replicas for its test window
metadata:
annotations:
hibernator.io/freeze-baseline: "true"

Lifecycle (an external test tool):

  1. Resource runs with baseline=1 (tracked by hibernator)
  2. The tool sets hibernator.io/freeze-baseline: "true" on the Deployment
  3. The tool sets KEDA paused-replicas: "2" — replicas scale to 2
  4. Hibernator’s discovery loop sees replicas=2 but skips baseline update
  5. The tool removes paused-replicas, replicas return to 1
  6. The tool removes freeze-baseline annotation
  7. Baseline remains 1 throughout — no corruption

Important:

  • Must be set before external scaling occurs
  • Must be removed after the test window ends
  • New resources discovered with this annotation still record their initial baseline (can’t freeze what hasn’t been stored yet)
  • If the annotation is never removed (e.g., CronJob failure), the baseline is permanently frozen. Monitor via standard CronJob alerting
  • Alerts for unexpected replica changes are suppressed while the annotation is active