Stop Hibernator in an emergency
Describes Hibernator chart 0.12.44
A stand-down makes Hibernator hand back everything it scaled down: it wakes the workloads, resumes the CronJobs, starts the databases and removes the satellites. After that, Hibernator scales nothing until the stand-down ends. This page tells you when to use it, how to give it, how to see that it finished and how to end it.
- When to use it
- Stand down in Manual Control
- Stand down through the API
- Stand down through Helm
- What the handback does
- See that the handback finished
- End the stand-down
- What a stand-down reports
When to use it
Section titled “When to use it”Stand Hibernator down when every workload must come back now, and Hibernator must not scale anything after that. For example:
- Hibernator scales workloads that it must leave alone, and you do not know the cause yet.
- An incident or a release needs every workload up until further notice.
- You want to uninstall Hibernator (Uninstall). Without a stand-down, the workloads that are hibernated at that time stay at zero replicas, because nothing wakes them.
A stand-down applies to the whole cluster. It has two sources, and either source alone holds it:
- An admin gives it in Manual Control or through the API, with an optional reason.
Hibernator keeps it in the
hibernator-stateConfigMap. It holds across a restart, an upgrade and a reinstall into the same namespace, until an admin ends it. - The configuration declares it with
config.global.enabled: false. Its actor isconfiguration. Only a change of the value totrueends it.
Dry run is not a stand-down. With config.operations.dryRun: true, the controller
logs each decision and changes nothing, so a hibernated workload stays at zero replicas.
Dry run freezes the cluster. A stand-down hands back also while dry run is on.
Stand down in Manual Control
Section titled “Stand down in Manual Control”Use this way when you can sign in to the UI as an admin.
- Sign in as an admin.
- Open Manual Control.
- In the Stand-down card, click Stand down. The dialog shows what the handback does now: how many workloads scale up, CronJobs resume and databases start, and whether satellites go.
- Optional: type a reason in Reason (optional). Use at most 50 characters. Do not put a URL in the reason, because the API refuses it.
- Click Stand down.
The card then shows who stood Hibernator down, and whether the handback runs or is finished. If Manual Control cannot read the status, the Stand down button is disabled. Use the API or Helm then.
Stand down through the API
Section titled “Stand down through the API”Use this way for a script, or when the UI does not load but the API answers. The API
needs a sign-in as an admin, and so it needs a working mail server
(Authentication). You also need kubectl and jq.
-
Forward a local port to the Hibernator service:
Terminal window kubectl port-forward -n hibernator svc/hibernator-app 8080:80 -
In a second terminal, request a one-time code for an admin address:
Terminal window curl -s -X POST http://127.0.0.1:8080/api/v1/auth/request-otp \-H 'Content-Type: application/json' -d '{"email": "admin@example.com"}' -
Get an access token with the code from the email:
Terminal window TOKEN=$(curl -s -X POST http://127.0.0.1:8080/api/v1/auth/verify-otp \-H 'Content-Type: application/json' \-d '{"email": "admin@example.com", "otp": "ABC234"}' | jq -r .token) -
Stand Hibernator down. The body with the reason is optional:
Terminal window curl -s -X POST http://127.0.0.1:8080/api/v1/stand-down \-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \-d '{"reason": "release freeze"}'
A 200 answer means that Hibernator stored the stand-down. The controller then starts the
handback. What a stand-down reports lists every refusal.
Stand down through Helm
Section titled “Stand down through Helm”Use this way when the UI or the sign-in does not work, or when your configuration lives in Git.
-
Find the chart version that runs now. The
CHARTcolumn shows it, for examplehibernator-X.Y.Z:Terminal window helm list -n hibernator -
In your values file, set
config.global.enabledtofalse:config:global:enabled: false -
Apply the values file with the same chart version:
Terminal window helm upgrade hibernator oci://registry.gitlab.com/cirriton/hibernator/charts/hibernator \-n hibernator -f my-values.yaml --version X.Y.Z
Helm changes the configuration ConfigMap, and no pod restarts. The controller reads the
change and starts the handback at once. The UI and the API show this stand-down after the
controller has seen it. Keep the value in your values file: the next helm upgrade with
true ends the stand-down.
What the handback does
Section titled “What the handback does”The handback is a wake, and the controller starts it at once:
- Deployments and StatefulSets go back to the replica count that Hibernator recorded for them. A workload recorded at 0 stays at 0.
- CronJobs that Hibernator suspended resume. CronJobs that their owner suspended stay suspended.
- Managed databases that Hibernator stopped start.
- Satellites go, and the service selectors are restored.
- Karpenter protection is released.
- A manual hibernation (Trigger Manual Hibernation) is cleared, as a wake clears it.
The workloads come up in your wake order (Wake order). A
hibernation that runs when the stand-down arrives stops. A step that fails is tried again
on every pass until everything is back. Until then the controller log has
Handback not finished, retrying on the next pass, and its pending field names what is
still missing. A controller that has just started hands back after its first license
check.
After the handback, Hibernator discovers nothing and scales nothing. It deletes nothing
and keeps its records. Nothing hibernates on schedule, and Trigger Manual Hibernation,
Wake Up Resources and Resume Schedule are disabled. The API refuses these actions
with 409 stood_down. Reads, configuration changes, external wake requests and the
stand-down endpoints work as usual.
With external wake, a stood-down cluster still answers the wake requests of its dependent clusters, and their wakes succeed: nothing it manages is down. As a dependent, it sends no wake requests, and its sessions on the hosting cluster expire.
See that the handback finished
Section titled “See that the handback finished”These show that the stand-down holds. They appear when it starts, and they do not change when the handback finishes:
- The red top-bar chip
Not scaling - stood down by <actor>, for every signed-in user. For the configuration it readsNot scaling - stood down by configuration. - The audit entry Stood Down.
- The Kubernetes Event
StoodDown. - The card
Stood Down by <actor>, if your notifications send it.
The handback itself shows as a wake. Its progress card reads Waking Up Resources and
Initiated by: <actor> (stand-down), and Wake Complete follows. When nothing is
hibernated, no wake card opens.
These show that the handback finished:
-
In Manual Control, the Stand-down card reads “The handback is finished: Hibernator woke everything it had scaled down.” While it runs, the card reads “The handback is running: Hibernator is waking everything it had scaled down.”
-
The dialog of the chip, Home and Manual Control say “It woke everything it had scaled down”. While the handback runs, they say “It is waking”.
-
GET /api/v1/hibernation/statusanswersstandDown.handbackDone: true. -
The controller log has the line
Handback finished: Hibernator has stopped scaling:Terminal window kubectl logs -n hibernator deployment/hibernator-controller | grep Handback
No card marks the end of the handback. The Wake Complete card can arrive before a database or a satellite is done.
End the stand-down
Section titled “End the stand-down”An admin ends an admin’s stand-down, and only a change of the configuration ends the stand-down from the configuration. When both sources hold, end each one.
An admin’s stand-down. In Manual Control:
- In the Stand-down card, click End stand-down. The dialog says what happens at the end.
- Click End stand-down.
Through the API, send DELETE /api/v1/stand-down with an admin’s token, as in
Stand down through the API.
The stand-down from the configuration.
- In your values file, set
config.global.enabledtotrue. - Run
helm upgradewith that file, as in Stand down through Helm.
The API refuses to end this stand-down with 409 stand_down_by_configuration, and Manual
Control shows no End stand-down button for it.
When the last cause ends, Hibernator scales again at once, and the schedule alone decides:
- Inside a hibernation window, Hibernator hibernates the cluster at once. The audit entry
of that hibernation names
<actor> (stand-down ended). - Outside a hibernation window, the cluster stays up. A manual hibernation from before the stand-down does not come back, because the handback cleared it. To hibernate the cluster, use Trigger Manual Hibernation.
While the other source or an unlicensed verdict still holds, Hibernator still scales nothing (License states).
What a stand-down reports
Section titled “What a stand-down reports”In the API, both endpoints are for admins only: 401 without a sign-in, 403 for
anyone else.
| Request | What it does | Refused with |
|---|---|---|
POST /api/v1/stand-down, optional body {"reason": "..."} |
gives an admin’s stand-down | 409 already_stood_down while an admin’s stand-down holds; 400 reason_contains_url for a reason that contains a URL |
DELETE /api/v1/stand-down |
ends the admin’s stand-down | 409 stand_down_by_configuration while only config.global.enabled: false holds one; 409 not_stood_down while none holds |
Both answer 200 with standDown: the stand-down that holds after the request, or
null. GET /api/v1/hibernation/status carries the same summary, with source, actor,
at, reason, configuration and handbackDone. The server-sent event
stand-down-update sends it to every signed-in user when it changes. While a stand-down
holds, an action that would scale answers 409 stood_down, and details.causes names
each cause that holds: stand_down, license.
Events on the release Namespace, in the release namespace:
kubectl get events -n hibernator --field-selector involvedObject.kind=Namespace| Event | Type | When |
|---|---|---|
StoodDown |
Warning | an admin or the configuration stands Hibernator down |
StandDownEnded |
Normal | a source of the stand-down ends; the message says whether Hibernator scales again |
Notifications go through the routing in Notifications. The event
key standDown sends both cards. Each destination takes it unless its events map sets
standDown or medium to false:
| Card | When |
|---|---|
Stood Down by <actor> |
an admin or the configuration stands Hibernator down |
Stand-down Ended by <actor> |
a source of the stand-down ends |
The actor is the admin’s address as your privacy settings show it, or configuration.
The handback’s outcome is the ordinary Wake Complete or Wake Failed card,
triggered by <actor> (stand-down).
Audit. The audit log records Stood Down (stand-down) and Stand-down Ended
(stand-down-ended), by the admin or by configuration, with the reason when the admin
gave one. A hibernation that starts when a stand-down ends is recorded as started by
<actor> (stand-down ended).
Log. The controller logs Stood down: Hibernator hands back everything it scaled, then stops scaling when a source begins, and Stand-down ended when one ends.