Skip to content

Stop Hibernator in an emergency

Describes Hibernator chart 0.12.44

A stand-down makes Hibernator hand back everything it scaled down: it wakes the workloads, resumes the CronJobs, starts the databases and removes the satellites. After that, Hibernator scales nothing until the stand-down ends. This page tells you when to use it, how to give it, how to see that it finished and how to end it.

Stand Hibernator down when every workload must come back now, and Hibernator must not scale anything after that. For example:

  • Hibernator scales workloads that it must leave alone, and you do not know the cause yet.
  • An incident or a release needs every workload up until further notice.
  • You want to uninstall Hibernator (Uninstall). Without a stand-down, the workloads that are hibernated at that time stay at zero replicas, because nothing wakes them.

A stand-down applies to the whole cluster. It has two sources, and either source alone holds it:

  • An admin gives it in Manual Control or through the API, with an optional reason. Hibernator keeps it in the hibernator-state ConfigMap. It holds across a restart, an upgrade and a reinstall into the same namespace, until an admin ends it.
  • The configuration declares it with config.global.enabled: false. Its actor is configuration. Only a change of the value to true ends it.

Dry run is not a stand-down. With config.operations.dryRun: true, the controller logs each decision and changes nothing, so a hibernated workload stays at zero replicas. Dry run freezes the cluster. A stand-down hands back also while dry run is on.

Use this way when you can sign in to the UI as an admin.

  1. Sign in as an admin.
  2. Open Manual Control.
  3. In the Stand-down card, click Stand down. The dialog shows what the handback does now: how many workloads scale up, CronJobs resume and databases start, and whether satellites go.
  4. Optional: type a reason in Reason (optional). Use at most 50 characters. Do not put a URL in the reason, because the API refuses it.
  5. Click Stand down.

The card then shows who stood Hibernator down, and whether the handback runs or is finished. If Manual Control cannot read the status, the Stand down button is disabled. Use the API or Helm then.

Use this way for a script, or when the UI does not load but the API answers. The API needs a sign-in as an admin, and so it needs a working mail server (Authentication). You also need kubectl and jq.

  1. Forward a local port to the Hibernator service:

    Terminal window
    kubectl port-forward -n hibernator svc/hibernator-app 8080:80
  2. In a second terminal, request a one-time code for an admin address:

    Terminal window
    curl -s -X POST http://127.0.0.1:8080/api/v1/auth/request-otp \
    -H 'Content-Type: application/json' -d '{"email": "admin@example.com"}'
  3. Get an access token with the code from the email:

    Terminal window
    TOKEN=$(curl -s -X POST http://127.0.0.1:8080/api/v1/auth/verify-otp \
    -H 'Content-Type: application/json' \
    -d '{"email": "admin@example.com", "otp": "ABC234"}' | jq -r .token)
  4. Stand Hibernator down. The body with the reason is optional:

    Terminal window
    curl -s -X POST http://127.0.0.1:8080/api/v1/stand-down \
    -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
    -d '{"reason": "release freeze"}'

A 200 answer means that Hibernator stored the stand-down. The controller then starts the handback. What a stand-down reports lists every refusal.

Use this way when the UI or the sign-in does not work, or when your configuration lives in Git.

  1. Find the chart version that runs now. The CHART column shows it, for example hibernator-X.Y.Z:

    Terminal window
    helm list -n hibernator
  2. In your values file, set config.global.enabled to false:

    config:
    global:
    enabled: false
  3. Apply the values file with the same chart version:

    Terminal window
    helm upgrade hibernator oci://registry.gitlab.com/cirriton/hibernator/charts/hibernator \
    -n hibernator -f my-values.yaml --version X.Y.Z

Helm changes the configuration ConfigMap, and no pod restarts. The controller reads the change and starts the handback at once. The UI and the API show this stand-down after the controller has seen it. Keep the value in your values file: the next helm upgrade with true ends the stand-down.

The handback is a wake, and the controller starts it at once:

  • Deployments and StatefulSets go back to the replica count that Hibernator recorded for them. A workload recorded at 0 stays at 0.
  • CronJobs that Hibernator suspended resume. CronJobs that their owner suspended stay suspended.
  • Managed databases that Hibernator stopped start.
  • Satellites go, and the service selectors are restored.
  • Karpenter protection is released.
  • A manual hibernation (Trigger Manual Hibernation) is cleared, as a wake clears it.

The workloads come up in your wake order (Wake order). A hibernation that runs when the stand-down arrives stops. A step that fails is tried again on every pass until everything is back. Until then the controller log has Handback not finished, retrying on the next pass, and its pending field names what is still missing. A controller that has just started hands back after its first license check.

After the handback, Hibernator discovers nothing and scales nothing. It deletes nothing and keeps its records. Nothing hibernates on schedule, and Trigger Manual Hibernation, Wake Up Resources and Resume Schedule are disabled. The API refuses these actions with 409 stood_down. Reads, configuration changes, external wake requests and the stand-down endpoints work as usual.

With external wake, a stood-down cluster still answers the wake requests of its dependent clusters, and their wakes succeed: nothing it manages is down. As a dependent, it sends no wake requests, and its sessions on the hosting cluster expire.

These show that the stand-down holds. They appear when it starts, and they do not change when the handback finishes:

  • The red top-bar chip Not scaling - stood down by <actor>, for every signed-in user. For the configuration it reads Not scaling - stood down by configuration.
  • The audit entry Stood Down.
  • The Kubernetes Event StoodDown.
  • The card Stood Down by <actor>, if your notifications send it.

The handback itself shows as a wake. Its progress card reads Waking Up Resources and Initiated by: <actor> (stand-down), and Wake Complete follows. When nothing is hibernated, no wake card opens.

These show that the handback finished:

  • In Manual Control, the Stand-down card reads “The handback is finished: Hibernator woke everything it had scaled down.” While it runs, the card reads “The handback is running: Hibernator is waking everything it had scaled down.”

  • The dialog of the chip, Home and Manual Control say “It woke everything it had scaled down”. While the handback runs, they say “It is waking”.

  • GET /api/v1/hibernation/status answers standDown.handbackDone: true.

  • The controller log has the line Handback finished: Hibernator has stopped scaling:

    Terminal window
    kubectl logs -n hibernator deployment/hibernator-controller | grep Handback

No card marks the end of the handback. The Wake Complete card can arrive before a database or a satellite is done.

An admin ends an admin’s stand-down, and only a change of the configuration ends the stand-down from the configuration. When both sources hold, end each one.

An admin’s stand-down. In Manual Control:

  1. In the Stand-down card, click End stand-down. The dialog says what happens at the end.
  2. Click End stand-down.

Through the API, send DELETE /api/v1/stand-down with an admin’s token, as in Stand down through the API.

The stand-down from the configuration.

  1. In your values file, set config.global.enabled to true.
  2. Run helm upgrade with that file, as in Stand down through Helm.

The API refuses to end this stand-down with 409 stand_down_by_configuration, and Manual Control shows no End stand-down button for it.

When the last cause ends, Hibernator scales again at once, and the schedule alone decides:

  • Inside a hibernation window, Hibernator hibernates the cluster at once. The audit entry of that hibernation names <actor> (stand-down ended).
  • Outside a hibernation window, the cluster stays up. A manual hibernation from before the stand-down does not come back, because the handback cleared it. To hibernate the cluster, use Trigger Manual Hibernation.

While the other source or an unlicensed verdict still holds, Hibernator still scales nothing (License states).

In the API, both endpoints are for admins only: 401 without a sign-in, 403 for anyone else.

Request What it does Refused with
POST /api/v1/stand-down, optional body {"reason": "..."} gives an admin’s stand-down 409 already_stood_down while an admin’s stand-down holds; 400 reason_contains_url for a reason that contains a URL
DELETE /api/v1/stand-down ends the admin’s stand-down 409 stand_down_by_configuration while only config.global.enabled: false holds one; 409 not_stood_down while none holds

Both answer 200 with standDown: the stand-down that holds after the request, or null. GET /api/v1/hibernation/status carries the same summary, with source, actor, at, reason, configuration and handbackDone. The server-sent event stand-down-update sends it to every signed-in user when it changes. While a stand-down holds, an action that would scale answers 409 stood_down, and details.causes names each cause that holds: stand_down, license.

Events on the release Namespace, in the release namespace:

Terminal window
kubectl get events -n hibernator --field-selector involvedObject.kind=Namespace
Event Type When
StoodDown Warning an admin or the configuration stands Hibernator down
StandDownEnded Normal a source of the stand-down ends; the message says whether Hibernator scales again

Notifications go through the routing in Notifications. The event key standDown sends both cards. Each destination takes it unless its events map sets standDown or medium to false:

Card When
Stood Down by <actor> an admin or the configuration stands Hibernator down
Stand-down Ended by <actor> a source of the stand-down ends

The actor is the admin’s address as your privacy settings show it, or configuration. The handback’s outcome is the ordinary Wake Complete or Wake Failed card, triggered by <actor> (stand-down).

Audit. The audit log records Stood Down (stand-down) and Stand-down Ended (stand-down-ended), by the admin or by configuration, with the reason when the admin gave one. A hibernation that starts when a stand-down ends is recorded as started by <actor> (stand-down ended).

Log. The controller logs Stood down: Hibernator hands back everything it scaled, then stops scaling when a source begins, and Stand-down ended when one ends.