Troubleshooting
Describes Hibernator chart 0.12.44
Hibernator’s transition events — wake and hibernation completions, wake barrier timeouts,
recovery catch-ups — are emitted against the Namespace object, so they land in the default
namespace and not in the namespace hibernator runs in:
kubectl get events -n default --sort-by=.lastTimestamp | grep -i hibernat# Controllerkubectl logs -n hibernator deployment/hibernator-controller -f
# APIkubectl logs -n hibernator deployment/hibernator-app -c api -f
# Frontendkubectl logs -n hibernator deployment/hibernator-app -c frontend -f
# Everythingkubectl logs -n hibernator -l app.kubernetes.io/name=hibernator --all-containers=true -fSet config.operations.logLevel: debug for more detail. The setting is read from the
ConfigMap at runtime, so it takes effect without a restart.
Inspect what was installed
Section titled “Inspect what was installed”# The rendered configurationkubectl get configmap -n hibernator hibernator-config -o yaml
# Secrets exist (names only, never the contents)kubectl get secret -n hibernator
# RBACkubectl get clusterrole,clusterrolebinding -l app.kubernetes.io/name=hibernatorkubectl get role,rolebinding -n hibernator
# Everything the release createdkubectl get all -n hibernator
# Wait for the controllerkubectl wait --for=condition=ready pod -l component=controller -n hibernator --timeout=60sCommon problems
Section titled “Common problems”The install fails on the schedule format
Section titled “The install fails on the schedule format”- at '/config/workingHours/schedule/monday': '8:00-20:00' does not match pattern '^(off|on|([0-1][0-9]|2[0-3]):[0-5][0-9]-([0-1][0-9]|2[0-3]):[0-5][0-9](,([0-1][0-9]|2[0-3]):[0-5][0-9]-([0-1][0-9]|2[0-3]):[0-5][0-9])*)$'Write each hour and each minute with two digits: 08:00-20:00, not 8:00-20:00.
schedule.md lists the values the chart refuses.
The install fails naming a required value
Section titled “The install fails naming a required value”config.clusterName is required. Set it in your values file (e.g., 'production-eu-central-1')secrets.jwtSecret is required - use: openssl rand -base64 32, or name your own Secret in secrets.existingSecretsecrets.internalApiSecret is required - use: openssl rand -base64 32, or name your own Secret in secrets.existingSecretThese values have no defaults. Set config.clusterName. Set the two secrets, or name
your own Secret in secrets.existingSecret: see
install.md.
Pods cannot pull images
Section titled “Pods cannot pull images”ImagePullBackOff, or Failed to pull image: unauthorized. The images are in a
private registry, so the release needs credentials:
helm upgrade hibernator oci://registry.gitlab.com/cirriton/hibernator/charts/hibernator -n hibernator \ --reuse-values \ --set imagePullSecrets.create=true \ --set imagePullSecrets.username=YOUR_USERNAME \ --set imagePullSecrets.password=YOUR_PASSWORDIf you manage the pull secret yourself, point imagePullSecrets.existingSecret at it
instead and leave create false.
Nobody receives a login code
Section titled “Nobody receives a login code”Login is email one-time-password only, so this means the install cannot send mail. Check, in order:
config.smtp.enabledistrueandconfig.smtp.hostis set. With SMTP off, the API rejects every login attempt withSMTP is disabled.secrets.smtp.userandsecrets.smtp.passwordare correct. Withsecrets.existingSecret, the keyssmtp-userandsmtp-passwordof your Secret are correct.- The address is allowed. If
config.auth.allowedEmailDomainsis non-empty, an address outside those domains is refused before any mail is attempted.
kubectl logs -n hibernator deployment/hibernator-app -c api | grep -i smtpA login code is refused with “too many attempts”
Section titled “A login code is refused with “too many attempts””The third wrong code for an address spends its outstanding code and the magic link that came with it, so even the right code is refused after that. Request a new code; that lifts the limit. Anyone who knows an address can spend its outstanding code this way, so a person who did not mistype may still see the message.
A pod stops at startup with “invalid configuration”
Section titled “A pod stops at startup with “invalid configuration””The API and the controller run the same validation at startup that a configuration
reload runs: working hours, targeting, resource rules and the URL readiness check. A
configuration a reload would refuse therefore stops the pod instead of booting. The
message says which check failed, for example invalid exclude pattern 0: cannot specify both pattern and exact: a targeting entry sets exact or pattern, not both.
kubectl logs -n hibernator deployment/hibernator-controller | grep -i "invalid configuration"Everyone on a domain was signed out at once
Section titled “Everyone on a domain was signed out at once”config.auth.allowedEmailDomains is re-read every time an access token is
issued, not only at the login screen. Removing a domain from the list therefore
signs its people out within config.auth.tokenExpiry – when their current
access token runs out and the refresh is refused – rather than letting them
carry on until their refresh token expires days later. The login screen tells
them the domain is not allowed, and lists the ones that are.
Add the domain back if the removal was a mistake; there is nothing else to clean up, because no session state is kept.
Passkeys fail after the passkey data was lost or restored
Section titled “Passkeys fail after the passkey data was lost or restored”Someone deleted the hibernator-passkeys ConfigMap, or restored it from a backup. Nobody
is locked out, because the one-time code always works. Read
Backup and restore. It tells what each case costs and
what to tell your users.
Permission denied in the controller log
Section titled “Permission denied in the controller log”pods is forbidden: User "system:serviceaccount:hibernator:hibernator-controller" cannot list resource "pods"The ClusterRole or its binding did not get created — usually because the installing account cannot create cluster-scoped RBAC.
kubectl get clusterrole hibernator-controllerkubectl get clusterrolebinding hibernator-controllerThe controller only ever has one replica
Section titled “The controller only ever has one replica”That is deliberate and cannot be changed. The controller runs exactly one replica with
a Recreate strategy so that two controllers can never reconcile the same resources at
once. replicaCount.app scales the API and frontend; nothing scales the controller.
The transition overlay is empty, or reports the wrong transition
Section titled “The transition overlay is empty, or reports the wrong transition”Check replicaCount.app first. The real-time event stream and the last-frame replay
behind the overlay both live in the memory of one API pod, and the controller posts
each event to the Service, which lands it on one pod. Above one replica a browser is
balanced onto a pod that may have seen none of the transition, so the overlay stays
empty – or onto one holding an earlier transition’s outcome, which opens its
completion dialog over the transition that is running.
kubectl get deploy -n hibernator hibernator-app -o jsonpath='{.spec.replicas}'The chart ships one replica, which is the supported setting. One-time codes and the sign-in rate limiter have the same limit; see Authentication.
Cost Explorer reports an access error
Section titled “Cost Explorer reports an access error”Failed to get daily costs from AWS: AccessDeniedException in the API log.
# Does the service account carry the role annotation?kubectl get sa -n hibernator hibernator-api -o yaml | grep eks.amazonaws.comThen check that the IAM role’s trust policy names that service account and that its
policy includes ce:GetCostAndUsage. See aws-integration.md.
Separately, the aws:eks:cluster-name cost allocation tag has to be activated in the
AWS Billing console under Cost Allocation Tags, or Cost Explorer returns nothing
for the filter regardless of permissions.
System Status Unavailable
Section titled “System Status Unavailable”The Home and Manual Control pages lead with a warning card instead of the status, and
the API answers 503 on the status endpoints:
kubectl port-forward -n hibernator svc/hibernator-app 8080:80curl -s http://127.0.0.1:8080/api/v1/public/hibernation-status{ "code": "managed_state_unavailable", "message": "The hibernator state could not be read; it recovers on its own once the read succeeds, no reload needed."}The instance refuses to say whether the cluster is down rather than guess. A guessed
“awake” would send visitors to applications that are still scaled to zero, so no
active is served at all. GET /wake/status, the dashboard status, POST /external-wake, POST /hibernation/trigger and POST /wake/up answer the same way,
and Manual Control refuses both hibernate and wake for as long as it holds — the refusal
comes before any write, so a refused trigger changes nothing. The controller is
unaffected: whatever the resources are doing, they keep doing it.
The four resource listings — GET /resources, GET /resources/{namespace},
GET /resources/ignored and GET /resources/stats — answer it too, but only for the
second cause below. They read the managed state in order to serve it, so a value that
will not decode leaves them nothing to list; a read that fails on the wire stays a
500 on all four. The Resources page shows the message GET /resources sent, in place
of the grids, so a resource_states that will not decode names its key on the page
rather than only in the log. The other three answer a caller of the API; the page
renders an empty ignored grid and no statistics for them, and says nothing.
There are two causes, and the message tells them apart:
-
The state could not be read — the wording above. The apiserver was unreachable, the read timed out, or RBAC refuses it. Nothing has to be repaired: the pages poll every 30 seconds and clear themselves on the first answer. If it does not clear, check the API log for the read failure and that the API’s service account may still read ConfigMaps in the release namespace.
-
A value cannot be decoded — the message instead names one key of the
hibernator-stateConfigMap:{"code": "managed_state_unavailable","message": "The \"resource_states\" value in the hibernator-state ConfigMap cannot be decoded, so every read returns the same value; an operator has to repair it before the state can be read again. Once it is repaired, the state reads again on the next poll, with no reload."}This one does not clear itself. Every poll reads the same value, so it holds until someone repairs it. The keys that can put this screen into it are
resource_statesandwake_override. Three other keys degrade instead, and never produce it:ignored_resources— a value there that will not decode still answersresources.excluded: 0on the dashboard status, on a200, and the pages carry on. It does refusePOST /external-wakefor named resources andGET /resources/ignored, though, with the same message naming that key. Neither refusal is visible on a page: the Resources page renders an empty ignored grid rather than the message. So an operator whose selective external wakes are all refused, or whose ignored list is empty while these pages read fine, should curl/api/v1/resources/ignoredand look atignored_resourcesas well.external_wake_sessions— a value there that will not decode still answersexternalWake: {enabled: true, activeSessions: 0}on the dashboard status, on a200, so the sessions in force are not listed while the rest of the response is unaffected.database_states— a value there that will not decode keeps the configured database count on the dashboard status and reports no statuses, on a200. It is logged at warn; the other two are logged at debug.
None of the three is visible on the wire: each answers exactly what a healthy cluster with no sessions, nothing excluded and no statuses yet would answer, so the API log is the only place the failure appears.
Look at the named key:
kubectl get configmap hibernator-state -n hibernator -o yamlEach holds JSON — resource_states, ignored_resources and external_wake_sessions
are arrays, wake_override an object, database_states an object keyed by database; a
truncated or hand-edited value is the usual cause. Repair it if you can read what it was
meant to be:
kubectl edit configmap hibernator-state -n hibernatorOtherwise clear the key to its empty value. That is safe but not free: an empty
resource_states is a cluster with nothing recorded as scaled down, so a cluster that
is currently hibernated loses the replica counts it would be woken back to, and those
workloads have to be scaled back by hand once. An empty wake_override only ends any
manual wake session that was in force. An empty external_wake_sessions ends every
inbound external wake session in force at once, so a dependent cluster that holds this
one awake has to ask again.
# nothing recorded as managedkubectl patch configmap hibernator-state -n hibernator --type merge -p '{"data":{"resource_states":"[]"}}'
# no manual wake in forcekubectl patch configmap hibernator-state -n hibernator --type merge -p '{"data":{"wake_override":""}}'
# no inbound external wake sessions in forcekubectl patch configmap hibernator-state -n hibernator --type merge -p '{"data":{"external_wake_sessions":"[]"}}'An absent hibernator-state is not this problem: a cluster that has never
hibernated has no ConfigMap yet, that is a valid empty state, and it answers normally.
Both pages recover on their own once the value reads again — the next poll, within 30 seconds, with no reload.
Nothing hibernates
Section titled “Nothing hibernates”Check in this order:
- Hibernator has not stopped scaling because of its license or a stand-down. A red
Not scalingchip in the top bar says it has. So does alicense verdictline in the controller log with the statehandback, or aStood downline. See Hibernator has stopped scaling. config.global.enabledistrue. The valuefalseis a stand-down.config.operations.dryRunisfalse. In dry-run the controller logs every decision and changes nothing — which is the recommended first-install setting, and easy to forget about afterwards.- The workload is not excluded:
config.targeting.namespaces.exclude,config.resourceRules.neverScale, or ahibernator.io/excludeannotation on the resource itself. See annotations.md. - It is genuinely outside working hours in
config.workingHours.timezone, and no wake session is holding the cluster up. An inbound external wake session keeps the whole cluster awake for as long as it lasts, whatever resources it names.
Hibernator has stopped scaling
Section titled “Hibernator has stopped scaling”A red Not scaling - unlicensed or Not scaling - license check failed chip in the top
bar, Hibernate Now, Wake and Resume Schedule disabled, the HibernatorNotActing alert,
or 409 unlicensed from the API: the license check does not let Hibernator act. It woke
everything it had scaled down and scales nothing until a check succeeds. An amber
License check failing or License expires chip is the warning before that, with the
date it would stop.
Settings → License names the reason. license.md has one entry per reason, from a missing license to an STS that cannot be reached, and what never to do to get out of it.
A red Not scaling - stood down by <actor> chip, or 409 stood_down from the API, is a
stand-down. Hibernator woke everything it had scaled down and scales nothing until the
stand-down ends. details.causes in the 409 names every cause that holds, so a
stand-down while unlicensed names both. To end a stand-down:
- An admin’s stand-down: click End stand-down in Manual Control as an admin, or send
DELETE /api/v1/stand-down. - A stand-down from the configuration (the actor is
configuration): setconfig.global.enabledtotrue. Then runhelm upgrade. The API refuses to end this stand-down with409 stand_down_by_configuration.
When the stand-down ends inside a hibernation window, Hibernator hibernates the cluster at once. The audit entry of that hibernation names who ended the stand-down. Outside a hibernation window the cluster stays up. A Hibernate Now from before the stand-down does not come back, because the handback cleared it. To hibernate the cluster again, use Trigger Manual Hibernation after the end. Stop Hibernator in an emergency describes the stand-down: the handback, how to see that it finished, and its API, events, cards and audit entries.
Clicking Wake showed no progress overlay
Section titled “Clicking Wake showed no progress overlay”The cluster was already running, so there was nothing to wake. Hibernator treats that click as a wake override renewal rather than a transition: it moves the override’s expiry forward and stops there. There is no progress overlay, no completion dialog, no Teams card and no mail, because nothing happened that somebody who was not watching needs to be told about. What you do get is a toast naming the new expiry, and an audit entry titled Wake Made No Changes.
This is deliberate. A wake that scaled nothing settles in about ten seconds, and recording that as a wake duration drags down the estimate every real wake’s overlay counts down against.
A wake whose first minutes are spent starting a managed database and waiting at its group is not this case: it has acted, so it is a real transition with an overlay and a duration of its own, even though no replica has changed yet.
A wake reports workloads as failed, still starting, or recovered
Section titled “A wake reports workloads as failed, still starting, or recovered”A wake’s verdict has three classes, and they mean different things. Reading them apart is the whole point of the split.
Still starting. The wake ended before hibernator had an answer about this
workload. Nothing is known to be wrong with it: it was inside its own budget and
still coming up when something else ended the wake. The wake is titled Wake Complete, the card and the mail are green, and the audit entry is a success
whose result reads incomplete — clean is about what the wake lost, complete is
about what it observed, and a wake can be both. There is nothing to do. The entry
names the cause, and each cause is somebody else’s business:
| Cause | What it means |
|---|---|
fail_open |
Hibernator’s whole-transition limit ran out. That alone is not a failure: a wake that fell open and lost nothing is green like any other still-starting wake, and one that did lose something is failed through its casualties. Check whether the workload came up a few minutes later; if it routinely does, its startupProbe is declaring less than it needs |
walked_past |
A wake tier was released without waiting for the group it depended on. The Wake Barrier Timeout notification for that tier is the one to read |
superseded |
A hibernation started while the wake was still running — an expiring wake session, or a schedule boundary. This entry and the Kubernetes event are the wake’s entire record |
not_observed |
The controller restarted and the cluster had meanwhile stopped heading where the wake was driving it, so nobody watched the rest |
Failed. Hibernator stopped waiting for this workload, and the reason says which
budget ran out or what could not recover. This is the class worth acting on: the
wake is titled Wake Complete — N workloads failed, the card and the mail are red,
and the audit entry’s result is failed. Each casualty carries a full sentence in
the activity log, on the card, in the mail and in the audit entry’s
workloads_failed_detail:
| Reason | Where to look |
|---|---|
| was still Unschedulable after 15 minutes | The scheduler could place a pod on no node for the whole of phase 1. Its PodScheduled condition says why — resource requests, node selectors, taints, or a node autoscaler that never added a node |
| was still Pending with no node after 15 minutes | A pod had no node and the scheduler had not marked it Unschedulable. Either no scheduler reported on it at all — a scheduler that is down or backlogged, or a schedulerName no scheduler serves — or its PodScheduled condition names another reason, such as SchedulingGated for a pod whose scheduling gates have not been removed, or SchedulerError. The condition says which |
| its init containers had not finished after 15 minutes | A pod had a node, and an init container was still running or had failed and was waiting to run again. The init container’s own log is the place to start |
| was still pulling its image after 15 minutes | A pod had a node and the kubelet had not yet created its container (or its current init container), with no pull failing — usually a slow pull, not a failing one: the node’s disk and the image’s size. The kubelet reports a volume that has not mounted, or pod networking not yet set up, the same way; kubectl describe pod events say which |
| had not started every replica after 15 minutes | No pod was in any of the four states above, and at least one replica had no started container — often because a pod was never created at all, whether or not its siblings started. kubectl describe the workload and its ReplicaSet: a quota or admission refusal shows there as FailedCreate |
| could not pull its image in 15 minutes | Your registry, or the path to it, or the pull secret |
| could not create its container (…) | The kubelet refused to create the container and the parenthesis is its own reason — a Secret key that is not there, a mount that cannot be made, a command that is not executable. You fix the manifest, not the cluster. Not terminal on sight: create the missing Secret while the wake runs and the pod is created on the next kubelet sync |
| its image reference is not a valid reference | The manifest. Nobody needs to check whether the registry is up |
| its image is not on the node and its pull policy forbids fetching it | imagePullPolicy: Never with no pre-loaded image. No pull is ever attempted, so no amount of waiting helps |
| did not finish starting within its declared 20 minutes | The application. It asked for that budget itself — see below |
| did not finish starting within the default 10 minutes | The application, judged by the fallback because it declares no startupProbe |
| did not finish starting within the 30-minute cap | The workload declares more than hibernator will wait. This is a fact about hibernator’s configuration, not a verdict on the workload |
| restarted twice before it was ready | The container is dying before it ever becomes ready — a crash, or its own startupProbe killing it |
| could not be read | Hibernator could not read the workload’s spec. Usually RBAC, or an apiserver under load |
| hibernator could not scale it up: … | Kubernetes refused the scale-up, and the message is the one it returned |
The five after 15 minutes sentences name what hibernator saw when phase 1 ran out (see below). Where the workload’s pods were in different states, the sentence names the pod furthest from a started container: Unschedulable, then Pending, then init containers, then a pull. A refused container and a pull still failing outrank all of them.
Recovered. Hibernator stopped waiting for this workload and then watched it come up
anyway. An observation outranks a verdict, so the write-off is withdrawn outright and the
workload is not a casualty: the wake is titled plain Wake Complete, the card and the
mail are green, and the audit entry is a success. Nothing about it escalates — a degraded wake
additionally reaches whatever destination takes this cluster’s high-priority events, and a
recovering one does not, so it is delivered exactly where a clean wake of its kind always was. A wake that lost nothing
and recovered two workloads reads exactly like a wake that recovered none, apart from saying
so — the clause in a card’s title is the escalation signal, and a positive one in that slot
would teach you that a clause no longer means bad news.
Each recovery carries one sentence, composed once by the controller and rendered word for word everywhere:
| Where | What you see |
|---|---|
| Teams card | A Recovered block between the casualty bullets and the still-starting sentence, naming the first three and then “and N more” |
| The same block, with the same cut, as a panel of its own in the HTML body and a heading of its own in the plain-text one | |
| Activity log | The casualty line is replaced in place — same slot — by the same sentence, in green |
| Audit entry | workloads_recovered and workloads_recovered_detail |
| Kubernetes event | A counted clause, “2 workloads recovered” |
| Prometheus | hibernator_workload_recoveries_total{reason} — see monitoring.md |
The card and the mail name at most three recoveries, so a green card does not turn into a wall of text on the morning one wake recovers two dozen workloads. The activity log and the audit entry keep every one. The casualty list is never cut: every workload a wake lost is named, however many there are.
The sentence has two parts. It starts with the reason of the write-off it lifted — the same words the casualty line had — and then says how the workload came up:
did not finish starting within its declared 20 minutes, then came up less than 1 minute after the restart
The reason tells you which kind of write-off was lifted: a first start that ran past its probe budget, an image pull that came back, a container that restarted twice before it was ready. The second part is one of these shapes:
| After the reason | What it means |
|---|---|
| came up less than 1 minute after the restart | A container restarted and that attempt came ready. The gap is measured from the restart, not from the write-off |
| came up less than 1 minute after hibernator stopped waiting | Nothing restarted. The application simply needed more seconds, or a phase-1 verdict’s registry came back |
| came up after hibernator stopped waiting | The verdict carried no instant to measure a gap from |
| came back up after hibernator stopped waiting | A workload that recovered, broke again and came back without ever being written off a second time. It states no gap, because every anchor the first one was measured from went with the write-off the lift cleared. The reason is still the one of the write-off that was lifted |
| any of the above, plus , and is not ready again | The lift still stands, and the workload is not ready right now |
Audit entries written before 0.11.147 carry recoveries with the second part alone in
workloads_recovered_detail. They keep the words they were written with.
A workload that recovers and then breaks again keeps its line rather than losing it: the sentence stays and gains “, and is not ready again”, in the casualty colour. If a second write-off is filed, the casualty line takes the slot back and the recovery in between is not recorded — the wake reports what it ended with.
A written-off workload gets up to five more minutes to come back
Section titled “A written-off workload gets up to five more minutes to come back”Hibernator does not end a wake on a write-off while a restart is in flight. A workload it has stopped waiting for, with a container visibly having another go at starting, is given up to five more minutes before the wake settles. That is what turns most would-be casualties into recoveries: without it the wake ended about ten seconds after the write-off, and a restart that came ready twenty seconds later missed its own correction.
There is nothing to configure and nothing to do. Four things worth knowing:
- It is evidence, not a longer deadline. Raising a
startupProbebudget to buy slack does not help here, because that budget is also what the kubelet kills on. What earns the extra time is a container that is actually running behind the verdict. - It is one band per workload per wake. A workload that keeps flapping does not keep earning five minutes, and a verdict that lifts and comes back gets none.
- A workload written off on its declared
startupProbebudget waits for the kubelet’s kill first. At the moment its budget runs out nothing has restarted yet: the kubelet kills the pod on its own probe a little later, and the restart is what the five minutes are for. So the wake does not settle until one probe period plus the probe’s timeouts after the budget —periodSeconds+timeoutSeconds+ the termination grace period, about 40 seconds with the defaults, and never more than the five minutes above. The kubelet counts the budget afresh for every attempt, so where a container restarted during that budget and its next attempt is the one still running, the wait is counted from where that attempt’s own budget runs out instead — still within the five minutes. The restart then gets whatever is left of the five minutes, counted from the write-off, and the workload is failed only if that attempt is not Ready when they end. A workload with nostartupProbe, or one held to the 30-minute cap, gets no such wait: nothing will kill its pod, so there is nothing to wait for. - A wake can therefore run a few minutes past its stated fail-open — see below.
The visible cost is a crashlooping workload delaying its wake’s card by up to five minutes, once. That is the price of the workloads that do come back.
How a workload declares how long it needs
Section titled “How a workload declares how long it needs”Each workload is judged against its own declaration, not against one number for the cluster. The deadline runs in two phases:
- Fifteen minutes for the cluster to produce a started container — node provisioning, scheduling, image pull. This is the cluster’s job, and it is a constant with no value to set.
- The budget the workload’s own
startupProbedeclares, which isinitialDelaySeconds+failureThreshold×periodSeconds, taken as the longest among the containers that declare one. The delay counts because the kubelet honours it before the first probe attempt. A sidecar that declares nothing is making no claim and neither raises nor lowers the budget. Capped at 30 minutes.
startupProbe: httpGet: { path: /healthz, port: 8080 } initialDelaySeconds: 0 periodSeconds: 10 failureThreshold: 120 # 0 + 10 x 120 = 20 minutesThat workload gets twenty minutes of phase 2, so fifteen plus twenty, and hibernator waits 35 minutes before writing it off. Raising the number in the manifest raises the budget: the declaration is made by the people who know, it already exists in most charts, and hibernator needs no annotation, no ConfigMap and no cooperation to read it.
A workload that declares no startupProbe gets config.controller.workloadSettleTimeout
for phase 2 — ten minutes by default. That value means exactly one thing: what a
workload that declares nothing gets. Ten minutes is a guess, and it is deliberately a
short one, because a workload that needs longer has a way to say so and a guess that
covered the slowest workload in the fleet would hold every wake open for it. A
workload whose spec cannot be read gets the same fallback rather than zero or the cap.
The whole transition has a fail-open behind those deadlines, and it is derived from the same declarations: the highest ceiling any of your workloads can reach, plus two minutes, never under fifteen. There is no value to set, and it cannot disagree with the per-workload deadlines it bounds, being derived from the same numbers.
A wake can legitimately run past that figure. While a written-off workload is inside the five-minute grace above, or waiting for the kubelet’s kill, the fail-open holds off with it — it fires at the later of the derived deadline and the latest live band’s or kill wait’s ending across every written-off workload, plus a 45-second settle margin. So the outside figure for a wake is the derived fail-open plus 3m45s, or plus 5m45s for a wake a controller restart resumed. The two move together deliberately: if the fail-open fired underneath a live grace, the wake would report “hibernator’s overall limit ran out” about a workload it had a real per-workload verdict for. Size a runbook off the outside figure, not off the derived one.
Why did I get two cards for one thing — or only one for two?
Section titled “Why did I get two cards for one thing — or only one for two?”Hibernator’s notification deduplicator suppresses a repeated message for
config.notifications.deduplication.windowSeconds (five minutes by default). What it
governs is deliberately narrow:
- Condition alerts are deduplicated. These are the things hibernator re-detects on every reconcile pass for as long as they are true — a workload that will not scale, an oversized state ConfigMap or Parameter Store, a state mismatch, an unexpected replica change, a schedule that will not parse, a config reload that failed. Without the window a stuck workload would produce a card per pass. This is what the window is for, and the value is worth raising if a persistent condition is still too loud.
- Reports of something that happened are exempt and always sent. A completed wake or hibernation, a catch-up for a boundary missed while the controller was down, a wake barrier timeout, a resource added or removed, an external wake started or expired, a schedule override, a database stopped. A transition reports its outcome once, so every time one of these was suppressed it was a defect: two wakes of the same cluster three minutes apart are two events, whatever they carry, and the second of them may be the one with a recovery to report.
So: two cards in quick succession for two real wakes is correct, and one card for a condition
that is still true is correct. External Wake Renewed is the one act still under the window,
because a requesting cluster re-asks on its own reconcile for as long as it holds the session.
A retried message is not a second event. The controller stamps every event it composes
with an identity, so a delivery the API retried is recognisably the same occurrence rather
than a second one and is suppressed — by the same window, which is the one thing the window
still does for a report. Setting windowSeconds very low therefore trades a quieter
condition alert for the chance of a retried wake card arriving twice.
Audit entries from before 0.11.91 read failed
Section titled “Audit entries from before 0.11.91 read failed”Audit entries written before 0.11.91 stamp a slow but successful wake as
result: failed, with workloads_not_ready and workloads_not_ready_detail in their
details. Hibernator then judged every workload by one flat deadline and fed every
workload short of ready into a single degraded flag, so a workload that declared
twenty minutes and came up in nineteen was recorded as a casualty of a failed wake.
Those entries are not migrated, deliberately. Rewriting an audit log to make
history look better is the one thing an audit log must not do. A query for
result = failed over a range that spans the upgrade will therefore return wakes that
were fine; the boundary is 0.11.91, and entries from it on use failed,
incomplete and success with the meanings above. The retired keys stay on the old
entries exactly as they were written.
Entries written before 0.11.147 also stamp result: failed on a wake whose overall
limit ran out with nothing lost: until then that alone made a wake failed. From 0.11.147
such a wake records incomplete at most, and its still-starting list says what had not
finished; with nothing still starting it records success.
A port-forward to the UI lands on config.externalUrl
Section titled “A port-forward to the UI lands on config.externalUrl”Expected. This instance has two own addresses, config.externalUrl and
config.internalUrl (which falls back to the external one when unset). A document
request is exempt from the redirect only when its host:port and its first path
segment both match the same one of them; every other document request gets
302 <config.externalUrl>#/?origin=<the URL you asked for>. That is what turns a
visitor who opened a hibernating application around before the login page boots, and a
port-forward address is indistinguishable from one: 127.0.0.1:8080 is not an address
this instance answers on. There is no exemption for localhost or an IP literal.
Because the two halves are matched per address rather than pooled, a request that
takes the host of one address and the path prefix of the other is not exempt. With
config.externalUrl: https://portal.example.com/hibernator/ and
config.internalUrl: https://hibernator-infra.example.com/, a document request to
https://hibernator-infra.example.com/hibernator/ is redirected: the host matches the
internal address, but that address’ first path segment is empty, so only / is exempt
on it. Reach the instance on the internal host at its own root, or on the external
host under its own prefix.
So this port-forward, from Verify the install, ends on
config.externalUrl:
kubectl port-forward -n hibernator svc/hibernator-app 8080:80# then http://127.0.0.1:8080 -> 302 to config.externalUrlWhat still works over the forward:
curl -s http://127.0.0.1:8080/api/v1/public/hibernation-status # the API, under /api/curl -s http://127.0.0.1:8080/healthz # the container probesTo open the UI, use config.externalUrl itself. If the instance is not routable yet,
the shortest way to a working forward is a hosts entry for the host in
config.externalUrl pointing at 127.0.0.1, forwarded on the port that URL names, so
the host:port you type is the one the instance answers on.
An external health check on / reports the frontend unhealthy
Section titled “An external health check on / reports the frontend unhealthy”The same rule. A load-balancer target group, an ingress healthCheck or an
uptime monitor pointed at / sends a Host that pairs with no own address of
this instance – a target IP, a Service name, the monitor’s own hostname – and
receives 302, not 200. A check that accepts only 200 marks every
target down.
Point such checks at /healthz, which the frontend answers 200 OK on both
:3000 and :3443 and which the rule exempts, along with /readyz. The chart’s
own container probes use it for this reason. There is no way to switch
the rule off: config.externalUrl is required, and the rule is what makes a
satellite address usable.
If tls.enabled is false, the certificate on :3443 comes from the internal CA of
Caddy. Configure an HTTPS check on that port so that it does not verify the certificate.
One related consequence of frontend.trustedProxies: the ?origin= the redirect
carries takes its scheme from X-Forwarded-Proto only when the peer is a trusted
proxy, and from the listener otherwise. With the default empty list and TLS
terminated at an ingress in front of the pod, an https visitor is sent back to
an http:// origin after the wake. Set frontend.trustedProxies to the ranges
your ingress originates from – which is what makes the API see real client
addresses too.