Skip to content

Troubleshooting

Describes Hibernator chart 0.12.44

Hibernator’s transition events — wake and hibernation completions, wake barrier timeouts, recovery catch-ups — are emitted against the Namespace object, so they land in the default namespace and not in the namespace hibernator runs in:

Terminal window
kubectl get events -n default --sort-by=.lastTimestamp | grep -i hibernat
Terminal window
# Controller
kubectl logs -n hibernator deployment/hibernator-controller -f
# API
kubectl logs -n hibernator deployment/hibernator-app -c api -f
# Frontend
kubectl logs -n hibernator deployment/hibernator-app -c frontend -f
# Everything
kubectl logs -n hibernator -l app.kubernetes.io/name=hibernator --all-containers=true -f

Set config.operations.logLevel: debug for more detail. The setting is read from the ConfigMap at runtime, so it takes effect without a restart.

Terminal window
# The rendered configuration
kubectl get configmap -n hibernator hibernator-config -o yaml
# Secrets exist (names only, never the contents)
kubectl get secret -n hibernator
# RBAC
kubectl get clusterrole,clusterrolebinding -l app.kubernetes.io/name=hibernator
kubectl get role,rolebinding -n hibernator
# Everything the release created
kubectl get all -n hibernator
# Wait for the controller
kubectl wait --for=condition=ready pod -l component=controller -n hibernator --timeout=60s
- at '/config/workingHours/schedule/monday': '8:00-20:00' does not match pattern '^(off|on|([0-1][0-9]|2[0-3]):[0-5][0-9]-([0-1][0-9]|2[0-3]):[0-5][0-9](,([0-1][0-9]|2[0-3]):[0-5][0-9]-([0-1][0-9]|2[0-3]):[0-5][0-9])*)$'

Write each hour and each minute with two digits: 08:00-20:00, not 8:00-20:00. schedule.md lists the values the chart refuses.

config.clusterName is required. Set it in your values file (e.g., 'production-eu-central-1')
secrets.jwtSecret is required - use: openssl rand -base64 32, or name your own Secret in secrets.existingSecret
secrets.internalApiSecret is required - use: openssl rand -base64 32, or name your own Secret in secrets.existingSecret

These values have no defaults. Set config.clusterName. Set the two secrets, or name your own Secret in secrets.existingSecret: see install.md.

ImagePullBackOff, or Failed to pull image: unauthorized. The images are in a private registry, so the release needs credentials:

Terminal window
helm upgrade hibernator oci://registry.gitlab.com/cirriton/hibernator/charts/hibernator -n hibernator \
--reuse-values \
--set imagePullSecrets.create=true \
--set imagePullSecrets.username=YOUR_USERNAME \
--set imagePullSecrets.password=YOUR_PASSWORD

If you manage the pull secret yourself, point imagePullSecrets.existingSecret at it instead and leave create false.

Login is email one-time-password only, so this means the install cannot send mail. Check, in order:

  1. config.smtp.enabled is true and config.smtp.host is set. With SMTP off, the API rejects every login attempt with SMTP is disabled.
  2. secrets.smtp.user and secrets.smtp.password are correct. With secrets.existingSecret, the keys smtp-user and smtp-password of your Secret are correct.
  3. The address is allowed. If config.auth.allowedEmailDomains is non-empty, an address outside those domains is refused before any mail is attempted.
Terminal window
kubectl logs -n hibernator deployment/hibernator-app -c api | grep -i smtp

A login code is refused with “too many attempts”

Section titled “A login code is refused with “too many attempts””

The third wrong code for an address spends its outstanding code and the magic link that came with it, so even the right code is refused after that. Request a new code; that lifts the limit. Anyone who knows an address can spend its outstanding code this way, so a person who did not mistype may still see the message.

A pod stops at startup with “invalid configuration”

Section titled “A pod stops at startup with “invalid configuration””

The API and the controller run the same validation at startup that a configuration reload runs: working hours, targeting, resource rules and the URL readiness check. A configuration a reload would refuse therefore stops the pod instead of booting. The message says which check failed, for example invalid exclude pattern 0: cannot specify both pattern and exact: a targeting entry sets exact or pattern, not both.

Terminal window
kubectl logs -n hibernator deployment/hibernator-controller | grep -i "invalid configuration"

Everyone on a domain was signed out at once

Section titled “Everyone on a domain was signed out at once”

config.auth.allowedEmailDomains is re-read every time an access token is issued, not only at the login screen. Removing a domain from the list therefore signs its people out within config.auth.tokenExpiry – when their current access token runs out and the refresh is refused – rather than letting them carry on until their refresh token expires days later. The login screen tells them the domain is not allowed, and lists the ones that are.

Add the domain back if the removal was a mistake; there is nothing else to clean up, because no session state is kept.

Passkeys fail after the passkey data was lost or restored

Section titled “Passkeys fail after the passkey data was lost or restored”

Someone deleted the hibernator-passkeys ConfigMap, or restored it from a backup. Nobody is locked out, because the one-time code always works. Read Backup and restore. It tells what each case costs and what to tell your users.

pods is forbidden: User "system:serviceaccount:hibernator:hibernator-controller" cannot list resource "pods"

The ClusterRole or its binding did not get created — usually because the installing account cannot create cluster-scoped RBAC.

Terminal window
kubectl get clusterrole hibernator-controller
kubectl get clusterrolebinding hibernator-controller

That is deliberate and cannot be changed. The controller runs exactly one replica with a Recreate strategy so that two controllers can never reconcile the same resources at once. replicaCount.app scales the API and frontend; nothing scales the controller.

The transition overlay is empty, or reports the wrong transition

Section titled “The transition overlay is empty, or reports the wrong transition”

Check replicaCount.app first. The real-time event stream and the last-frame replay behind the overlay both live in the memory of one API pod, and the controller posts each event to the Service, which lands it on one pod. Above one replica a browser is balanced onto a pod that may have seen none of the transition, so the overlay stays empty – or onto one holding an earlier transition’s outcome, which opens its completion dialog over the transition that is running.

Terminal window
kubectl get deploy -n hibernator hibernator-app -o jsonpath='{.spec.replicas}'

The chart ships one replica, which is the supported setting. One-time codes and the sign-in rate limiter have the same limit; see Authentication.

Failed to get daily costs from AWS: AccessDeniedException in the API log.

Terminal window
# Does the service account carry the role annotation?
kubectl get sa -n hibernator hibernator-api -o yaml | grep eks.amazonaws.com

Then check that the IAM role’s trust policy names that service account and that its policy includes ce:GetCostAndUsage. See aws-integration.md. Separately, the aws:eks:cluster-name cost allocation tag has to be activated in the AWS Billing console under Cost Allocation Tags, or Cost Explorer returns nothing for the filter regardless of permissions.

The Home and Manual Control pages lead with a warning card instead of the status, and the API answers 503 on the status endpoints:

Terminal window
kubectl port-forward -n hibernator svc/hibernator-app 8080:80
curl -s http://127.0.0.1:8080/api/v1/public/hibernation-status
{
"code": "managed_state_unavailable",
"message": "The hibernator state could not be read; it recovers on its own once the read succeeds, no reload needed."
}

The instance refuses to say whether the cluster is down rather than guess. A guessed “awake” would send visitors to applications that are still scaled to zero, so no active is served at all. GET /wake/status, the dashboard status, POST /external-wake, POST /hibernation/trigger and POST /wake/up answer the same way, and Manual Control refuses both hibernate and wake for as long as it holds — the refusal comes before any write, so a refused trigger changes nothing. The controller is unaffected: whatever the resources are doing, they keep doing it.

The four resource listings — GET /resources, GET /resources/{namespace}, GET /resources/ignored and GET /resources/stats — answer it too, but only for the second cause below. They read the managed state in order to serve it, so a value that will not decode leaves them nothing to list; a read that fails on the wire stays a 500 on all four. The Resources page shows the message GET /resources sent, in place of the grids, so a resource_states that will not decode names its key on the page rather than only in the log. The other three answer a caller of the API; the page renders an empty ignored grid and no statistics for them, and says nothing.

There are two causes, and the message tells them apart:

  • The state could not be read — the wording above. The apiserver was unreachable, the read timed out, or RBAC refuses it. Nothing has to be repaired: the pages poll every 30 seconds and clear themselves on the first answer. If it does not clear, check the API log for the read failure and that the API’s service account may still read ConfigMaps in the release namespace.

  • A value cannot be decoded — the message instead names one key of the hibernator-state ConfigMap:

    {
    "code": "managed_state_unavailable",
    "message": "The \"resource_states\" value in the hibernator-state ConfigMap cannot be decoded, so every read returns the same value; an operator has to repair it before the state can be read again. Once it is repaired, the state reads again on the next poll, with no reload."
    }

    This one does not clear itself. Every poll reads the same value, so it holds until someone repairs it. The keys that can put this screen into it are resource_states and wake_override. Three other keys degrade instead, and never produce it:

    • ignored_resources — a value there that will not decode still answers resources.excluded: 0 on the dashboard status, on a 200, and the pages carry on. It does refuse POST /external-wake for named resources and GET /resources/ignored, though, with the same message naming that key. Neither refusal is visible on a page: the Resources page renders an empty ignored grid rather than the message. So an operator whose selective external wakes are all refused, or whose ignored list is empty while these pages read fine, should curl /api/v1/resources/ignored and look at ignored_resources as well.
    • external_wake_sessions — a value there that will not decode still answers externalWake: {enabled: true, activeSessions: 0} on the dashboard status, on a 200, so the sessions in force are not listed while the rest of the response is unaffected.
    • database_states — a value there that will not decode keeps the configured database count on the dashboard status and reports no statuses, on a 200. It is logged at warn; the other two are logged at debug.

    None of the three is visible on the wire: each answers exactly what a healthy cluster with no sessions, nothing excluded and no statuses yet would answer, so the API log is the only place the failure appears.

Look at the named key:

Terminal window
kubectl get configmap hibernator-state -n hibernator -o yaml

Each holds JSON — resource_states, ignored_resources and external_wake_sessions are arrays, wake_override an object, database_states an object keyed by database; a truncated or hand-edited value is the usual cause. Repair it if you can read what it was meant to be:

Terminal window
kubectl edit configmap hibernator-state -n hibernator

Otherwise clear the key to its empty value. That is safe but not free: an empty resource_states is a cluster with nothing recorded as scaled down, so a cluster that is currently hibernated loses the replica counts it would be woken back to, and those workloads have to be scaled back by hand once. An empty wake_override only ends any manual wake session that was in force. An empty external_wake_sessions ends every inbound external wake session in force at once, so a dependent cluster that holds this one awake has to ask again.

Terminal window
# nothing recorded as managed
kubectl patch configmap hibernator-state -n hibernator --type merge -p '{"data":{"resource_states":"[]"}}'
# no manual wake in force
kubectl patch configmap hibernator-state -n hibernator --type merge -p '{"data":{"wake_override":""}}'
# no inbound external wake sessions in force
kubectl patch configmap hibernator-state -n hibernator --type merge -p '{"data":{"external_wake_sessions":"[]"}}'

An absent hibernator-state is not this problem: a cluster that has never hibernated has no ConfigMap yet, that is a valid empty state, and it answers normally.

Both pages recover on their own once the value reads again — the next poll, within 30 seconds, with no reload.

Check in this order:

  1. Hibernator has not stopped scaling because of its license or a stand-down. A red Not scaling chip in the top bar says it has. So does a license verdict line in the controller log with the state handback, or a Stood down line. See Hibernator has stopped scaling.
  2. config.global.enabled is true. The value false is a stand-down.
  3. config.operations.dryRun is false. In dry-run the controller logs every decision and changes nothing — which is the recommended first-install setting, and easy to forget about afterwards.
  4. The workload is not excluded: config.targeting.namespaces.exclude, config.resourceRules.neverScale, or a hibernator.io/exclude annotation on the resource itself. See annotations.md.
  5. It is genuinely outside working hours in config.workingHours.timezone, and no wake session is holding the cluster up. An inbound external wake session keeps the whole cluster awake for as long as it lasts, whatever resources it names.

A red Not scaling - unlicensed or Not scaling - license check failed chip in the top bar, Hibernate Now, Wake and Resume Schedule disabled, the HibernatorNotActing alert, or 409 unlicensed from the API: the license check does not let Hibernator act. It woke everything it had scaled down and scales nothing until a check succeeds. An amber License check failing or License expires chip is the warning before that, with the date it would stop.

Settings → License names the reason. license.md has one entry per reason, from a missing license to an STS that cannot be reached, and what never to do to get out of it.

A red Not scaling - stood down by <actor> chip, or 409 stood_down from the API, is a stand-down. Hibernator woke everything it had scaled down and scales nothing until the stand-down ends. details.causes in the 409 names every cause that holds, so a stand-down while unlicensed names both. To end a stand-down:

  • An admin’s stand-down: click End stand-down in Manual Control as an admin, or send DELETE /api/v1/stand-down.
  • A stand-down from the configuration (the actor is configuration): set config.global.enabled to true. Then run helm upgrade. The API refuses to end this stand-down with 409 stand_down_by_configuration.

When the stand-down ends inside a hibernation window, Hibernator hibernates the cluster at once. The audit entry of that hibernation names who ended the stand-down. Outside a hibernation window the cluster stays up. A Hibernate Now from before the stand-down does not come back, because the handback cleared it. To hibernate the cluster again, use Trigger Manual Hibernation after the end. Stop Hibernator in an emergency describes the stand-down: the handback, how to see that it finished, and its API, events, cards and audit entries.

The cluster was already running, so there was nothing to wake. Hibernator treats that click as a wake override renewal rather than a transition: it moves the override’s expiry forward and stops there. There is no progress overlay, no completion dialog, no Teams card and no mail, because nothing happened that somebody who was not watching needs to be told about. What you do get is a toast naming the new expiry, and an audit entry titled Wake Made No Changes.

This is deliberate. A wake that scaled nothing settles in about ten seconds, and recording that as a wake duration drags down the estimate every real wake’s overlay counts down against.

A wake whose first minutes are spent starting a managed database and waiting at its group is not this case: it has acted, so it is a real transition with an overlay and a duration of its own, even though no replica has changed yet.

A wake reports workloads as failed, still starting, or recovered

Section titled “A wake reports workloads as failed, still starting, or recovered”

A wake’s verdict has three classes, and they mean different things. Reading them apart is the whole point of the split.

Still starting. The wake ended before hibernator had an answer about this workload. Nothing is known to be wrong with it: it was inside its own budget and still coming up when something else ended the wake. The wake is titled Wake Complete, the card and the mail are green, and the audit entry is a success whose result reads incomplete — clean is about what the wake lost, complete is about what it observed, and a wake can be both. There is nothing to do. The entry names the cause, and each cause is somebody else’s business:

Cause What it means
fail_open Hibernator’s whole-transition limit ran out. That alone is not a failure: a wake that fell open and lost nothing is green like any other still-starting wake, and one that did lose something is failed through its casualties. Check whether the workload came up a few minutes later; if it routinely does, its startupProbe is declaring less than it needs
walked_past A wake tier was released without waiting for the group it depended on. The Wake Barrier Timeout notification for that tier is the one to read
superseded A hibernation started while the wake was still running — an expiring wake session, or a schedule boundary. This entry and the Kubernetes event are the wake’s entire record
not_observed The controller restarted and the cluster had meanwhile stopped heading where the wake was driving it, so nobody watched the rest

Failed. Hibernator stopped waiting for this workload, and the reason says which budget ran out or what could not recover. This is the class worth acting on: the wake is titled Wake Complete — N workloads failed, the card and the mail are red, and the audit entry’s result is failed. Each casualty carries a full sentence in the activity log, on the card, in the mail and in the audit entry’s workloads_failed_detail:

Reason Where to look
was still Unschedulable after 15 minutes The scheduler could place a pod on no node for the whole of phase 1. Its PodScheduled condition says why — resource requests, node selectors, taints, or a node autoscaler that never added a node
was still Pending with no node after 15 minutes A pod had no node and the scheduler had not marked it Unschedulable. Either no scheduler reported on it at all — a scheduler that is down or backlogged, or a schedulerName no scheduler serves — or its PodScheduled condition names another reason, such as SchedulingGated for a pod whose scheduling gates have not been removed, or SchedulerError. The condition says which
its init containers had not finished after 15 minutes A pod had a node, and an init container was still running or had failed and was waiting to run again. The init container’s own log is the place to start
was still pulling its image after 15 minutes A pod had a node and the kubelet had not yet created its container (or its current init container), with no pull failing — usually a slow pull, not a failing one: the node’s disk and the image’s size. The kubelet reports a volume that has not mounted, or pod networking not yet set up, the same way; kubectl describe pod events say which
had not started every replica after 15 minutes No pod was in any of the four states above, and at least one replica had no started container — often because a pod was never created at all, whether or not its siblings started. kubectl describe the workload and its ReplicaSet: a quota or admission refusal shows there as FailedCreate
could not pull its image in 15 minutes Your registry, or the path to it, or the pull secret
could not create its container (…) The kubelet refused to create the container and the parenthesis is its own reason — a Secret key that is not there, a mount that cannot be made, a command that is not executable. You fix the manifest, not the cluster. Not terminal on sight: create the missing Secret while the wake runs and the pod is created on the next kubelet sync
its image reference is not a valid reference The manifest. Nobody needs to check whether the registry is up
its image is not on the node and its pull policy forbids fetching it imagePullPolicy: Never with no pre-loaded image. No pull is ever attempted, so no amount of waiting helps
did not finish starting within its declared 20 minutes The application. It asked for that budget itself — see below
did not finish starting within the default 10 minutes The application, judged by the fallback because it declares no startupProbe
did not finish starting within the 30-minute cap The workload declares more than hibernator will wait. This is a fact about hibernator’s configuration, not a verdict on the workload
restarted twice before it was ready The container is dying before it ever becomes ready — a crash, or its own startupProbe killing it
could not be read Hibernator could not read the workload’s spec. Usually RBAC, or an apiserver under load
hibernator could not scale it up: … Kubernetes refused the scale-up, and the message is the one it returned

The five after 15 minutes sentences name what hibernator saw when phase 1 ran out (see below). Where the workload’s pods were in different states, the sentence names the pod furthest from a started container: Unschedulable, then Pending, then init containers, then a pull. A refused container and a pull still failing outrank all of them.

Recovered. Hibernator stopped waiting for this workload and then watched it come up anyway. An observation outranks a verdict, so the write-off is withdrawn outright and the workload is not a casualty: the wake is titled plain Wake Complete, the card and the mail are green, and the audit entry is a success. Nothing about it escalates — a degraded wake additionally reaches whatever destination takes this cluster’s high-priority events, and a recovering one does not, so it is delivered exactly where a clean wake of its kind always was. A wake that lost nothing and recovered two workloads reads exactly like a wake that recovered none, apart from saying so — the clause in a card’s title is the escalation signal, and a positive one in that slot would teach you that a clause no longer means bad news.

Each recovery carries one sentence, composed once by the controller and rendered word for word everywhere:

Where What you see
Teams card A Recovered block between the casualty bullets and the still-starting sentence, naming the first three and then “and N more”
Email The same block, with the same cut, as a panel of its own in the HTML body and a heading of its own in the plain-text one
Activity log The casualty line is replaced in place — same slot — by the same sentence, in green
Audit entry workloads_recovered and workloads_recovered_detail
Kubernetes event A counted clause, “2 workloads recovered”
Prometheus hibernator_workload_recoveries_total{reason} — see monitoring.md

The card and the mail name at most three recoveries, so a green card does not turn into a wall of text on the morning one wake recovers two dozen workloads. The activity log and the audit entry keep every one. The casualty list is never cut: every workload a wake lost is named, however many there are.

The sentence has two parts. It starts with the reason of the write-off it lifted — the same words the casualty line had — and then says how the workload came up:

did not finish starting within its declared 20 minutes, then came up less than 1 minute after the restart

The reason tells you which kind of write-off was lifted: a first start that ran past its probe budget, an image pull that came back, a container that restarted twice before it was ready. The second part is one of these shapes:

After the reason What it means
came up less than 1 minute after the restart A container restarted and that attempt came ready. The gap is measured from the restart, not from the write-off
came up less than 1 minute after hibernator stopped waiting Nothing restarted. The application simply needed more seconds, or a phase-1 verdict’s registry came back
came up after hibernator stopped waiting The verdict carried no instant to measure a gap from
came back up after hibernator stopped waiting A workload that recovered, broke again and came back without ever being written off a second time. It states no gap, because every anchor the first one was measured from went with the write-off the lift cleared. The reason is still the one of the write-off that was lifted
any of the above, plus , and is not ready again The lift still stands, and the workload is not ready right now

Audit entries written before 0.11.147 carry recoveries with the second part alone in workloads_recovered_detail. They keep the words they were written with.

A workload that recovers and then breaks again keeps its line rather than losing it: the sentence stays and gains “, and is not ready again”, in the casualty colour. If a second write-off is filed, the casualty line takes the slot back and the recovery in between is not recorded — the wake reports what it ended with.

A written-off workload gets up to five more minutes to come back

Section titled “A written-off workload gets up to five more minutes to come back”

Hibernator does not end a wake on a write-off while a restart is in flight. A workload it has stopped waiting for, with a container visibly having another go at starting, is given up to five more minutes before the wake settles. That is what turns most would-be casualties into recoveries: without it the wake ended about ten seconds after the write-off, and a restart that came ready twenty seconds later missed its own correction.

There is nothing to configure and nothing to do. Four things worth knowing:

  • It is evidence, not a longer deadline. Raising a startupProbe budget to buy slack does not help here, because that budget is also what the kubelet kills on. What earns the extra time is a container that is actually running behind the verdict.
  • It is one band per workload per wake. A workload that keeps flapping does not keep earning five minutes, and a verdict that lifts and comes back gets none.
  • A workload written off on its declared startupProbe budget waits for the kubelet’s kill first. At the moment its budget runs out nothing has restarted yet: the kubelet kills the pod on its own probe a little later, and the restart is what the five minutes are for. So the wake does not settle until one probe period plus the probe’s timeouts after the budget — periodSeconds + timeoutSeconds + the termination grace period, about 40 seconds with the defaults, and never more than the five minutes above. The kubelet counts the budget afresh for every attempt, so where a container restarted during that budget and its next attempt is the one still running, the wait is counted from where that attempt’s own budget runs out instead — still within the five minutes. The restart then gets whatever is left of the five minutes, counted from the write-off, and the workload is failed only if that attempt is not Ready when they end. A workload with no startupProbe, or one held to the 30-minute cap, gets no such wait: nothing will kill its pod, so there is nothing to wait for.
  • A wake can therefore run a few minutes past its stated fail-open — see below.

The visible cost is a crashlooping workload delaying its wake’s card by up to five minutes, once. That is the price of the workloads that do come back.

Each workload is judged against its own declaration, not against one number for the cluster. The deadline runs in two phases:

  1. Fifteen minutes for the cluster to produce a started container — node provisioning, scheduling, image pull. This is the cluster’s job, and it is a constant with no value to set.
  2. The budget the workload’s own startupProbe declares, which is initialDelaySeconds + failureThreshold × periodSeconds, taken as the longest among the containers that declare one. The delay counts because the kubelet honours it before the first probe attempt. A sidecar that declares nothing is making no claim and neither raises nor lowers the budget. Capped at 30 minutes.
startupProbe:
httpGet: { path: /healthz, port: 8080 }
initialDelaySeconds: 0
periodSeconds: 10
failureThreshold: 120 # 0 + 10 x 120 = 20 minutes

That workload gets twenty minutes of phase 2, so fifteen plus twenty, and hibernator waits 35 minutes before writing it off. Raising the number in the manifest raises the budget: the declaration is made by the people who know, it already exists in most charts, and hibernator needs no annotation, no ConfigMap and no cooperation to read it.

A workload that declares no startupProbe gets config.controller.workloadSettleTimeout for phase 2 — ten minutes by default. That value means exactly one thing: what a workload that declares nothing gets. Ten minutes is a guess, and it is deliberately a short one, because a workload that needs longer has a way to say so and a guess that covered the slowest workload in the fleet would hold every wake open for it. A workload whose spec cannot be read gets the same fallback rather than zero or the cap.

The whole transition has a fail-open behind those deadlines, and it is derived from the same declarations: the highest ceiling any of your workloads can reach, plus two minutes, never under fifteen. There is no value to set, and it cannot disagree with the per-workload deadlines it bounds, being derived from the same numbers.

A wake can legitimately run past that figure. While a written-off workload is inside the five-minute grace above, or waiting for the kubelet’s kill, the fail-open holds off with it — it fires at the later of the derived deadline and the latest live band’s or kill wait’s ending across every written-off workload, plus a 45-second settle margin. So the outside figure for a wake is the derived fail-open plus 3m45s, or plus 5m45s for a wake a controller restart resumed. The two move together deliberately: if the fail-open fired underneath a live grace, the wake would report “hibernator’s overall limit ran out” about a workload it had a real per-workload verdict for. Size a runbook off the outside figure, not off the derived one.

Why did I get two cards for one thing — or only one for two?

Section titled “Why did I get two cards for one thing — or only one for two?”

Hibernator’s notification deduplicator suppresses a repeated message for config.notifications.deduplication.windowSeconds (five minutes by default). What it governs is deliberately narrow:

  • Condition alerts are deduplicated. These are the things hibernator re-detects on every reconcile pass for as long as they are true — a workload that will not scale, an oversized state ConfigMap or Parameter Store, a state mismatch, an unexpected replica change, a schedule that will not parse, a config reload that failed. Without the window a stuck workload would produce a card per pass. This is what the window is for, and the value is worth raising if a persistent condition is still too loud.
  • Reports of something that happened are exempt and always sent. A completed wake or hibernation, a catch-up for a boundary missed while the controller was down, a wake barrier timeout, a resource added or removed, an external wake started or expired, a schedule override, a database stopped. A transition reports its outcome once, so every time one of these was suppressed it was a defect: two wakes of the same cluster three minutes apart are two events, whatever they carry, and the second of them may be the one with a recovery to report.

So: two cards in quick succession for two real wakes is correct, and one card for a condition that is still true is correct. External Wake Renewed is the one act still under the window, because a requesting cluster re-asks on its own reconcile for as long as it holds the session.

A retried message is not a second event. The controller stamps every event it composes with an identity, so a delivery the API retried is recognisably the same occurrence rather than a second one and is suppressed — by the same window, which is the one thing the window still does for a report. Setting windowSeconds very low therefore trades a quieter condition alert for the chance of a retried wake card arriving twice.

Audit entries from before 0.11.91 read failed

Section titled “Audit entries from before 0.11.91 read failed”

Audit entries written before 0.11.91 stamp a slow but successful wake as result: failed, with workloads_not_ready and workloads_not_ready_detail in their details. Hibernator then judged every workload by one flat deadline and fed every workload short of ready into a single degraded flag, so a workload that declared twenty minutes and came up in nineteen was recorded as a casualty of a failed wake.

Those entries are not migrated, deliberately. Rewriting an audit log to make history look better is the one thing an audit log must not do. A query for result = failed over a range that spans the upgrade will therefore return wakes that were fine; the boundary is 0.11.91, and entries from it on use failed, incomplete and success with the meanings above. The retired keys stay on the old entries exactly as they were written.

Entries written before 0.11.147 also stamp result: failed on a wake whose overall limit ran out with nothing lost: until then that alone made a wake failed. From 0.11.147 such a wake records incomplete at most, and its still-starting list says what had not finished; with nothing still starting it records success.

A port-forward to the UI lands on config.externalUrl

Section titled “A port-forward to the UI lands on config.externalUrl”

Expected. This instance has two own addresses, config.externalUrl and config.internalUrl (which falls back to the external one when unset). A document request is exempt from the redirect only when its host:port and its first path segment both match the same one of them; every other document request gets 302 <config.externalUrl>#/?origin=<the URL you asked for>. That is what turns a visitor who opened a hibernating application around before the login page boots, and a port-forward address is indistinguishable from one: 127.0.0.1:8080 is not an address this instance answers on. There is no exemption for localhost or an IP literal.

Because the two halves are matched per address rather than pooled, a request that takes the host of one address and the path prefix of the other is not exempt. With config.externalUrl: https://portal.example.com/hibernator/ and config.internalUrl: https://hibernator-infra.example.com/, a document request to https://hibernator-infra.example.com/hibernator/ is redirected: the host matches the internal address, but that address’ first path segment is empty, so only / is exempt on it. Reach the instance on the internal host at its own root, or on the external host under its own prefix.

So this port-forward, from Verify the install, ends on config.externalUrl:

Terminal window
kubectl port-forward -n hibernator svc/hibernator-app 8080:80
# then http://127.0.0.1:8080 -> 302 to config.externalUrl

What still works over the forward:

Terminal window
curl -s http://127.0.0.1:8080/api/v1/public/hibernation-status # the API, under /api/
curl -s http://127.0.0.1:8080/healthz # the container probes

To open the UI, use config.externalUrl itself. If the instance is not routable yet, the shortest way to a working forward is a hosts entry for the host in config.externalUrl pointing at 127.0.0.1, forwarded on the port that URL names, so the host:port you type is the one the instance answers on.

An external health check on / reports the frontend unhealthy

Section titled “An external health check on / reports the frontend unhealthy”

The same rule. A load-balancer target group, an ingress healthCheck or an uptime monitor pointed at / sends a Host that pairs with no own address of this instance – a target IP, a Service name, the monitor’s own hostname – and receives 302, not 200. A check that accepts only 200 marks every target down.

Point such checks at /healthz, which the frontend answers 200 OK on both :3000 and :3443 and which the rule exempts, along with /readyz. The chart’s own container probes use it for this reason. There is no way to switch the rule off: config.externalUrl is required, and the rule is what makes a satellite address usable.

If tls.enabled is false, the certificate on :3443 comes from the internal CA of Caddy. Configure an HTTPS check on that port so that it does not verify the certificate.

One related consequence of frontend.trustedProxies: the ?origin= the redirect carries takes its scheme from X-Forwarded-Proto only when the peer is a trusted proxy, and from the listener otherwise. With the default empty list and TLS terminated at an ingress in front of the pod, an https visitor is sent back to an http:// origin after the wake. Set frontend.trustedProxies to the ranges your ingress originates from – which is what makes the API see real client addresses too.