Skip to content

External wake

Describes Hibernator chart 0.12.44

Multi-cluster wake: one Hibernator asks another to wake resources on its cluster, so a frontend cluster can bring up the backend it depends on. Requests are HMAC-signed over HTTPS, and every cluster entry states exactly what the caller is allowed to wake.

Every pair has a requesting cluster and a receiving cluster.

  • The receiving cluster sets config.externalWake.enabled: true and names itself with config.externalWake.localClusterName.
  • The receiving cluster lists the callers it accepts under config.externalWake.clusters, together with the resources each caller may wake.
  • The requesting cluster lists what it depends on under config.externalDependencies, together with the endpoint to reach.
  • The requesting cluster signs each request with its config.clusterName. The receiving cluster’s clusters entry uses that name.

Both sides need the same HMAC secret for the pair. The chart refuses a key under config.externalWake that it does not read, such as dependencies.

Receiving cluster — authorization policy:

config:
externalWake:
enabled: true
localClusterName: "cluster-b"
clusters:
- name: "cluster-a"
secretRef:
name: cluster-hmac
key: shared-key
deployments:
- namespace: "app"
resources: ["backend-service"]

Requesting cluster — external dependencies:

config:
clusterName: "cluster-a"
externalDependencies:
- cluster: "cluster-b"
endpoint: "https://hibernator.cluster-b.example.com"
secretRef:
name: cluster-hmac
key: shared-key
deployments:
- namespace: "app"
resources: ["backend-service"]

A user triggering a wake on the requesting cluster automatically sends the signed request to every dependent cluster. Nothing else has to happen by hand.

You can use per-cluster secrets (a different secret for each cluster) or shared secrets (the same secret for several clusters). Both are supported.

Pattern matching for blue/green deployments

Section titled “Pattern matching for blue/green deployments”

Use glob patterns (*) in cluster names to support blue/green deployments without config changes:

# Cluster A values.yaml
externalWakeSecrets:
create: true
secrets:
- name: team-ab-hmac
key: shared-key
value: "ABC123..." # openssl rand -base64 32
config:
externalWake:
enabled: true
clusters:
- name: "team-b-*" # Matches team-b-01, team-b-02, etc.
secretRef:
name: team-ab-hmac
key: shared-key
# Cluster B values.yaml
externalWakeSecrets:
create: true
secrets:
- name: team-ab-hmac
key: shared-key
value: "ABC123..." # SAME VALUE as cluster A (recommended) or use different secrets per direction
config:
externalWake:
enabled: true
clusters:
- name: "team-a-*" # Matches team-a-01, team-a-02, etc.
secretRef:
name: team-ab-hmac
key: shared-key

For security isolation with many dependent clusters, use different secrets:

externalWakeSecrets:
create: true
secrets:
- name: cluster-a-hmac
key: shared-key
value: "ABC123..." # Unique secret for cluster A
- name: cluster-b-hmac
key: shared-key
value: "XYZ789..." # Different secret for cluster B
config:
externalWake:
enabled: true
clusters:
- name: "cluster-a-prod"
secretRef:
name: cluster-a-hmac
key: shared-key
- name: "cluster-b-prod"
secretRef:
name: cluster-b-hmac
key: shared-key

values.schema.json models config.externalWake.clusters[] and rejects any key it does not know, before a single template renders. The complete entry is:

Key Type Required
name string (glob patterns allowed) yes
secretRef.name / secretRef.key string yes
allowWakeAll boolean no
deployments / statefulSets / cronJobs list of { namespace, resources }, both required no

This matters because the ConfigMap template renders a cluster entry key by key and discards everything else. Before the schema modelled the entry, deploymnets: rendered successfully into a cluster that authorized nothing:

Terminal window
$ helm template ... --set 'config.externalWake.clusters[0].deploymnets[0].namespace=app'
Error: values don't meet the specifications of the schema(s) in the following chart(s):
hibernator:
- at '/config/externalWake/clusters/0': additional properties 'deploymnets' not allowed

For production, prefer manual secret creation or external secret managers (Vault, Sealed Secrets, External Secrets Operator):

Terminal window
# Generate HMAC secret
HMAC_SECRET=$(openssl rand -base64 32)
# Create secret on hosting cluster
kubectl create secret generic team-ab-hmac \
--namespace hibernator \
--from-literal=shared-key="$HMAC_SECRET"

Glob patterns in cluster names (team-b-*) match blue/green deployments without a configuration change on either side.

External wake sessions expire on their own when their duration runs out. An inbound session counts as an up reason for the whole cluster for as long as it holds, whatever resources it names — including for managed databases.

A dependent cluster asking again for resources it already holds renews the session it has rather than opening a second one: the hosting cluster moves that session’s expiresAt to the newly granted time and answers with the same requestID. So a second Wake click extends the window instead of spending another of maxConcurrentWakes, and the active-sessions list shows one session per resource set per dependent cluster rather than one per click. A renewal is granted under the same maxWakeDuration and working-hours caps as a first request, and it never shortens a session: asking for an hour while four hours are still held leaves the four hours in place. Asking for a different set of resources is a different session, and so is the first request after a session has expired.

A renewal notifies under its own headline, External Wake Renewed, so a Teams card or a mail subject says which of the two happened without being read to the end. It is enabled by the same externalWakeStarted toggle as a first grant.

A request answered 503 with managed_state_unavailable means the hosting cluster could not read its own managed state. Nothing was woken and no session was opened, so the dependent cluster should retry rather than treat its resources as unmanaged. A resource the host can read but does not manage is a different answer: 403 with reason not_found.

Retrying is the right client behaviour either way, but it only clears one of the two causes. The message says which: an unreachable apiserver or a read that timed out is worded as a read that failed and may well succeed on the next attempt, while a value of hibernator-state that will not decode is read again identically every time. That message names the key holding it — resource_states, ignored_resources or wake_override — and no amount of retrying clears it; an operator on the hosting cluster has to repair that value first, after which the next request is answered normally. A dependent cluster that retries forever on the second one is waiting for something that is not going to happen on its own, so the message is worth logging where somebody will see it.

A session belongs to the cluster named in the request body’s requestor.cluster: its quota is counted against that cluster, and only that cluster may read, renew or cancel it. So the body must name the cluster the HMAC signature authenticated. A request whose body names another cluster is answered 403 with code requestor_cluster_mismatch, and nothing is written: no session, no quota use, no audit entry. Hibernator’s own client always sends the same name in both.

To force-clear sessions, for a test or after a mistake:

Terminal window
kubectl get configmap hibernator-state -n hibernator -o json | \
jq '.data.external_wake_sessions = "[]"' | kubectl apply -f -

The controller picks the change up on its next reconcile, within about 30 seconds, and enforces the current schedule again.