Skip to content

License

Describes Hibernator chart 0.12.44

Hibernator acts only in a place its license names. This page covers everything about the license check: where it works, what it needs from AWS, how to install and replace a license, what users see, and what to do when the check does not say licensed. The stand-down, the other cause that stops Hibernator scaling, has its own page: Stop Hibernator in an emergency.

Hibernator contains a license check. § 7 of the LICENSE file that ships in this chart describes it, and your grant, the written permission from cirriton that you accepted, refers to it. To read the file, unpack the chart with helm pull oci://registry.gitlab.com/cirriton/hibernator/charts/hibernator --untar.

A license is one file, HL-XXXXXXXX.license, signed by cirriton. It names a licensee, the places it covers (its scopes, such as aws:111122223333 for an AWS account) and the last day it is valid. One license covers every place of one grant.

The controller proves the place it runs in, then compares that place with the license. It checks at start, every hour, and at once when the license or the configuration changes. Each check must finish within 30 seconds. The check runs inside the cluster: it sends no data to cirriton or to anyone else.

A check ends in one of three verdicts:

  • licensed: the place is proven and the license names it.
  • unlicensed: a definite no. There is no license, it is invalid or expired, the platform cannot be licensed, or the license does not name the place.
  • cannot tell: an error, such as an AWS STS that does not answer. It is never unlicensed. Hibernator keeps acting until 14 days after the last licensed check (grace).

On unlicensed, Hibernator first wakes everything it scaled down, then stops scaling. It never crashes, exits or refuses to start because of the license, and it deletes nothing.

The controller picks the platform from the providerID of the Node it runs on.

Platform When Scope it proves
AWS The controller’s Node has an aws:// providerID: EKS, EKS on Fargate, and clusters on EC2 that run the AWS cloud provider aws:<account>, the account of the AWS role the controller gets (AWS access)

On a kind cluster the check proves the place kind. No license for a licensee names it, so Hibernator does not scale on kind.

Any other platform is unlicensed with reason unsupported_platform, and no license can change that. This includes EKS hybrid nodes (eks-hybrid://), k3s, AKS, GKE, and a Node with no providerID. A Node whose cloud provider has not initialized it yet (it carries the node.cloudprovider.kubernetes.io/uninitialized taint) is cannot tell until it is.

The AWS account is the unit, not the cluster: a license that names an account covers every cluster in it, and a cluster rebuilt in the same account needs no new license.

On AWS the controller proves its account with sts:GetCallerIdentity, using the AWS role it gets from one of these three sources, tried in this order:

  1. IRSA (recommended): annotate the controller’s ServiceAccount with a role, as in aws-integration.md. It works on EC2 nodes and on Fargate.
  2. EKS Pod Identity: an association between a role and the controller’s ServiceAccount (hibernator-controller by default) in the release namespace.
  3. The node role, through instance metadata. Nothing to configure, but the pod must reach instance metadata: with IMDSv2 and a hop limit of 1, it cannot. Fargate has no node role.

Access keys are not accepted. A controller whose only AWS credentials are access keys gets the reason no_aws_role. If a source is configured but fails, the next one is tried, and the evidence on Settings → License names the source that proved the account.

No IAM permission is needed. GetCallerIdentity needs none, and no policy can deny it. If you use no other AWS feature of Hibernator, a role with no policy attached is enough; the controller role from aws-integration.md works too.

STS must be reachable from the controller pod. The check calls the regional STS endpoint, https://sts.<region>.amazonaws.com. The region is AWS_REGION or AWS_DEFAULT_REGION when one is set (IRSA sets both), otherwise config.awsCosts.region.

  • Without config.proxyURL, the call goes direct, or through HTTPS_PROXY when the controller’s environment sets it.
  • With config.proxyURL set, every STS call goes through that proxy, the IRSA token exchange included. NO_PROXY is not consulted, and there is no direct fallback. Instance metadata and the EKS Pod Identity agent never go through the proxy.
  • If a proxy inspects TLS with its own CA, put the CA in a PEM file and point AWS_CA_BUNDLE at it. The bundle replaces the system roots, so it must hold every root your path to AWS needs. It applies to Hibernator’s other AWS calls too:
extraEnvVars:
- name: AWS_CA_BUNDLE
value: /etc/hibernator/ca/bundle.pem
extraVolumes:
- name: aws-ca
configMap:
name: corporate-ca
extraVolumeMounts:
- name: aws-ca
mountPath: /etc/hibernator/ca
readOnly: true

The chart value license takes the license file, whole and unchanged. Use one of these:

  • --set-file, the simplest:

    Terminal window
    helm upgrade --install hibernator oci://registry.gitlab.com/cirriton/hibernator/charts/hibernator \
    -n hibernator -f my-values.yaml --set-file license=HL-XXXXXXXX.license
  • A values block:

    license: |
    -----BEGIN HIBERNATOR LICENSE-----
    <the lines of your license file>
    -----END HIBERNATOR LICENSE-----
  • Flux, inline: the same block under spec.values of the HelmRelease.

  • Flux, valuesFrom: a Secret whose key values.yaml holds the block above:

    spec:
    valuesFrom:
    - kind: Secret
    name: hibernator-license-values
    valuesKey: values.yaml
  • A ConfigMap you create yourself, and leave license empty:

    Terminal window
    kubectl create configmap hibernator-license -n hibernator \
    --from-file=license=HL-XXXXXXXX.license

    Never do both: when license is set, the chart creates the ConfigMap hibernator-license, and Helm fails if one it does not own exists.

Whitespace does not matter. Indentation, CRLF line ends, blank lines and re-wrapped lines around and inside the armor all read the same, so no Helm path or editor can break the file. The chart’s schema checks that the value runs from the -----BEGIN HIBERNATOR LICENSE----- line to the -----END HIBERNATOR LICENSE----- line, so a truncated or wrong paste fails helm install or helm upgrade. It does not check the signature: the controller does, and Settings → License shows what it found.

Only the controller reads the ConfigMap. It is not mounted, and a change restarts no pod.

The file is the license’s text and cirriton’s SSH signature, base64-encoded inside the armor. It is not encrypted. To read what it names:

Terminal window
sed '1d;$d' HL-XXXXXXXX.license | base64 -d

The output starts with the license:

format: "hibernator-license/v1"
id: "HL-XXXXXXXX"
issued: "2026-09-29"
expires: "2999-12-31"
licensee: "Example Licensee GmbH"
scopes:
- "aws:111122223333"

and ends with the -----BEGIN SSH SIGNATURE----- block. expires is the last day the license is valid, through 23:59:59 UTC. A license for a grant without an end date expires on 2999-12-31. Check that scopes names every place you install in: an AWS account as aws: and its 12 digits.

Put the new file in the license value (or in your ConfigMap) and run helm upgrade. The controller watches the ConfigMap hibernator-license, so the check runs at once and no pod restarts.

An installation holds exactly one license: the one in that ConfigMap. Hibernator never compares two licenses, and a newer file never takes over on its own. When your grant changes, cirriton issues a replacement license with the full list of places, and you install it.

Without a license the check says unlicensed, reason no_license. An emptied or deleted ConfigMap counts as no license. Hibernator then wakes everything it scaled down and stops scaling (see License states). A fresh install without a license scales nothing from the start.

The UI and the API stay up, and the check keeps running: installing a license is enough to make Hibernator scale again, with no restart.

A license names your accounts. If your rules keep such data out of plain Git:

  • Flux: the valuesFrom Secret from Install a license, encrypted with SOPS.
  • Argo CD: the license: block in a values file encrypted with SOPS, read through the helm-secrets plugin (secrets:// in valueFiles).
  • Helm: --set-file license=… with a file kept outside the repository.

The chart does not read the license from a Secret. The ConfigMap is readable by anyone who can read ConfigMaps in the release namespace.

The verdict says what a check found. The license state says what Hibernator does because of it:

State Hibernator When
licensed acts as usual the last check was licensed, or cannot tell within a day of the last licensed check
grace acts as usual, and warns with a stop date cannot tell for more than a day, and grace has not ended
handback wakes everything it scaled down unlicensed, or cannot tell with no grace left
inert does nothing but check the license after the handback

The UI and the API treat handback and inert alike: Hibernator has stopped scaling.

Grace. A cannot tell verdict counts as licensed for 14 days after the last licensed check, as long as the license still names the place that check proved. Grace never runs past the end of the license’s expiry day. For the first day nobody is told anything: a check that succeeds at least once a day is no news. After that day the state turns grace, and every surface below shows the stop date. A place whose check has never succeeded has no grace: it stops at once. A definite unlicensed never gets grace. The 14 days and the quiet day are fixed: no setting, and no license, changes them.

Expiry. A license is valid through 23:59:59 UTC of its expires day. From 30 days before, the chip, the alert, an event and a notification warn with the date, and another event and notification follow 3 days before. After that day the check says unlicensed, reason expired, without grace.

Stopped scaling is what handback and inert mean together. The handback is a wake, shown and reported like one, started by “license check”:

  • Deployments and StatefulSets go back to the replica count Hibernator recorded for them. A workload recorded at 0 stays at 0.
  • CronJobs that Hibernator suspended resume. CronJobs their owner suspended stay suspended.
  • Managed databases that Hibernator stopped start.
  • Satellites go, and service selectors are restored.
  • Karpenter protection is released.
  • A manual hibernation (Hibernate Now) is cleared, as Wake clears it.

The handback runs even with config.operations.dryRun: true, because waking is the harmless direction. A step that fails is retried on every pass until everything is back, and the controller log names what is still pending. From the start of the handback, the API refuses Hibernate Now, Wake and Resume Schedule.

Once the handback is done, Hibernator discovers nothing and scales nothing: nothing hibernates on schedule. It keeps its records, deletes nothing, and keeps checking. The first licensed verdict lets it scale again, with no restart. Then the schedule alone decides: inside a hibernation window the cluster hibernates, and outside it the cluster stays up.

A stand-down is the second cause of the same handback: an admin gives it in Manual Control, or config.global.enabled: false declares it. One handback serves both causes, and each cause alone keeps Hibernator from scaling. A licensed verdict during a stand-down therefore changes nothing until the stand-down ends, and the end of a stand-down while unlicensed changes nothing either. Stop Hibernator in an emergency tells how to give and end one, and what it reports.

With external wake, a cluster that has stopped scaling still answers the wake requests of its dependent clusters, and their wakes succeed: nothing it manages is down. As a dependent, it sends no wake requests, and its sessions on the hosting cluster expire.

In the web UI, a chip in the top bar shows the license:

Chip Who sees it When
Licensed to <licensee> (neutral) admins licensed, and the license ends in more than 30 days
License check failing - stops <date> (amber) everyone signed in grace
License expires - stops <date> (amber) everyone signed in the license ends within 30 days
Not scaling - unlicensed (red) everyone signed in stopped scaling on a definite no
Not scaling - license check failed (red) everyone signed in stopped scaling because the check could not tell, and no grace was left

A click opens a dialog with the state, the reason, what it means and the stop date. Users who are not admins never see the licensee, the place or the evidence, and the login page shows nothing about the license. The Help page explains the chip and grace.

Settings → License is the admin’s full record: the state, the verdict, the reason and its cause, the licensee, the license ID, its scopes, issue and expiry dates, the place seen (the scope the check proved, also when the license does not name it), the evidence of how it was proven, when the check that stored the record ran (the last licensed check, or the first one that saw the current verdict), since when the verdict holds, the grace end, and the fingerprint of the key that signed the license. The evidence on AWS names the caller, the role source and the STS endpoint, and says whether an endpoint override or a CA bundle was set:

arn:aws:sts::111122223333:assumed-role/hibernator/1700000000 through web identity (IRSA); STS endpoint https://sts.eu-central-1.amazonaws.com; no endpoint override; no CA bundle

When a new license would help, the section shows three steps, How to get a license, with the place seen and a Copy button.

While Hibernator has stopped scaling, Hibernate Now, Wake and Resume Schedule are disabled with a tooltip, and Home and Manual Control say so. During the handback, the transition card reads “Waking Up Resources”, with “Why: the license check stopped Hibernator: <reason>” and “license check” as the initiator.

In the API, the summary license: {state, reason, graceEndsAt, expiresAt} is part of GET /api/v1/version, GET /api/v1/public/hibernation-status and GET /api/v1/public/dashboard-status. These endpoints need no sign-in, so the summary never carries the licensee, the place or the evidence. It is absent until the first check has stored a verdict. expiresAt is the end of the expiry day: 00:00 UTC of the day after.

  • GET /api/v1/license is the full record, for admins only: 401 without a sign-in, 403 for anyone else, and 404 with code no_license_verdict before the first check.
  • The server-sent event license-update carries the summary to every signed-in user when it changes.
  • While Hibernator has stopped scaling, POST /api/v1/hibernation/trigger, POST /api/v1/hibernation/resume, POST /api/v1/wake/up and POST /api/v1/wake/sleep answer 409 Conflict with code unlicensed and the reason in details.reason. While a stand-down holds as well, the code is stood_down. details.causes names each cause that holds: license, stand_down. Reads, configuration changes, the external wake endpoint and the stand-down endpoints work as usual.

GET /api/v1/public/dashboard-status answers only while config.dashboardApi.enabled is true. It is the status endpoint for an external dashboard.

Alerts. With metrics.prometheusRule.enabled: true the PrometheusRule carries three license alerts, all warning: a license issue stops scaling and raises costs, it never takes a workload down. Each one can be turned off under metrics.prometheusRule.alerts, see monitoring.md.

Alert Fires when
HibernatorLicenseUnverified the check cannot tell and Hibernator still acts, from a day after the last licensed check; the description names the stop date
HibernatorNotActing Hibernator has not been acting for 5 minutes
HibernatorLicenseExpiring the license expires within 30 days

Metrics. Only the controller exposes them (/metrics on port 9090):

Metric Meaning
hibernator_license_verdict{verdict, reason} 1 for the verdict (licensed, unlicensed or cannot_tell) and reason of the latest check that decided
hibernator_license_acting 1 while Hibernator acts (licensed or grace); 0 when it has stopped scaling, and before the first verdict
hibernator_license_grace_ends_timestamp_seconds when grace ends or ended; 0 without a grace record
hibernator_license_expires_timestamp_seconds the end of the license’s expiry day; 0 without a valid license

Events. Kubernetes Events on the release Namespace, in the release namespace:

Terminal window
kubectl get events -n hibernator --field-selector involvedObject.kind=Namespace
Event Type When
LicenseUnverified Warning grace begins, with the stop date
NotActing Warning Hibernator stops scaling and the handback starts
Licensed Normal licensed again, after grace or after Hibernator stopped scaling
LicenseExpiring Warning the license expires within 30 days, and again within 3

Notifications go through the Teams and email routing in Notifications. Each destination takes all four unless its events map turns them off, one by one or with medium: false:

Event key Card When
licenseUnverified License Check Failing grace begins, and 3 days before the stop date
licenseNotActing Stopped Scaling Hibernator stops scaling and the handback starts
licenseRestored License Check Succeeded licensed again, after grace or after Hibernator stopped scaling
licenseExpiring License Expiring 30 days and 3 days before the license expires

Each one goes out once, from the check that crosses its edge, not once a day. The handback’s outcome is the ordinary Wake Complete or Wake Failed card, triggered by “license check”. No card or mail carries the licensee.

Log and audit. The controller logs one license verdict line per change of the verdict, its reason or the state, with the licensee, the place, the evidence, the reason and the grace end: INFO while the state is licensed, WARN otherwise. The audit log records license-verdict (by “license check”) for each change of the state, and of the verdict or its reason while the state is not licensed; the handback is license-handback.

Start with Settings → License: the reason, its cause and the place seen answer most questions. Without the UI:

Terminal window
kubectl logs -n hibernator deployment/hibernator-controller | grep 'license verdict'

For a new license or any question about one, write to the contact named in your grant. Send the place seen with a license request.

Never edit or delete the key license_verdict in the ConfigMap hibernator-state by hand. It holds the record grace is counted from: without it, a check that cannot tell stops Hibernator at once. Changing secrets.internalApiSecret has the same effect while the check cannot tell, because the record depends on it.

no_license: the ConfigMap hibernator-license is missing, or its key license is empty. Set the license value, or create the ConfigMap; see Install a license. Check the namespace: the controller reads the ConfigMap in its own namespace only.

invalid_license: the file is not a valid license. Its signature does not match (the file was changed after signing), the armor is broken, there is text outside it or a second license in it, or a field is missing or malformed. The cause names the rule. Install the file exactly as you received it; whitespace is never the problem.

untrusted_key: the license is intact, but this release does not know the key that signed it. The cause names the key’s fingerprint. If cirriton signed the license, a newer release trusts the key: send the fingerprint to your contact, who tells you the lowest release that does, and upgrade.

expired: the license’s expires day has passed (UTC). There is no grace. Install the renewed license.

unsupported_platform: the controller runs on a Node whose providerID belongs to no supported platform, or has none. The cause names the prefix. No license helps; run the controller on an AWS Node.

refused: the platform refused to confirm the place; the cause names why. No license helps.

scope_not_named: the check proved the place, and the license does not name it. The place seen shows which place that is, for example aws:111122223333. Check that you installed the right license; otherwise request one that names the place, with the place seen.

Each of these is cannot tell: Hibernator keeps acting until grace ends, or stops at once where no check has ever succeeded. Fix it before the stop date: the next check, at the latest within an hour, clears it. To check at once, change the license value or restart the controller; a restart keeps the grace it had.

no_aws_role: no source gave the controller an AWS role. Either none is configured (no IRSA annotation, no Pod Identity association, and instance metadata out of reach), or every configured source failed and the cause names each failure. Access keys are not accepted. See AWS access.

invalid_aws_settings: the AWS settings cannot be used: AWS_CA_BUNDLE points at a file that is missing or holds no certificate, config.proxyURL does not parse, or the AWS environment variables are malformed. The cause names which.

no_aws_region: the controller has no region: neither AWS_REGION nor AWS_DEFAULT_REGION is set, and config.awsCosts.region is empty. Set one of them.

aws_partition: the account is in an AWS partition other than aws, such as China or GovCloud. They are not supported.

sts_denied: STS answered 403. Check the role’s credentials and their expiry, and that the controller’s clock is right.

sts_error: STS answered with an error other than 403, a 5xx included, or answered the IRSA token exchange with an error other than a denial. A 5xx is AWS’s side; for any other answer, the cause has STS’s error code.

sts_unreachable: no answer from STS, or from the proxy. Allow the controller to reach sts.<region>.amazonaws.com, directly or through the proxy; check config.proxyURL, the proxy’s allow-list, DNS and network policies. A TLS error from an inspecting proxy lands here too: see AWS_CA_BUNDLE under AWS access.

malformed_response: STS answered 200 with a body that names no account, which points at something between the controller and STS rewriting the answer.

deadline: the check did not finish within 30 seconds: a slow proxy, STS, or instance metadata.

own_node_unknown: the controller does not know its Node, because the environment variable NODE_NAME is missing. The chart sets it from the downward API; check that nothing rewrites the controller Deployment.

node_read_failed: the controller could not read its own Node, or list the Nodes of the cluster. Check that the ClusterRole the chart creates still grants get and list on nodes.

node_uninitialized: the controller’s Node has no providerID yet and carries the node.cloudprovider.kubernetes.io/uninitialized taint. This clears once the cloud provider initializes the Node.

platform_error: the platform failed without naming a cause. The cause in Settings → License has the error.

0.12.0 is the first release that acts on the license check. Before you upgrade:

  1. Get your license, and check that it names every place you install in.
  2. Set license in your values, or create the ConfigMap hibernator-license. Charts before 0.12.0 accept the value and ignore it, so you can add it ahead of the upgrade.
  3. On AWS, make sure the controller gets an AWS role and can reach STS; see AWS access.

Then upgrade as usual. Without a license, the upgraded controller wakes everything it scaled down and stops scaling. When license is empty, the notes helm install and helm upgrade print say so in one line; you can ignore it if you created the ConfigMap yourself.

After the upgrade, Settings → License shows the state Licensed, and the controller log has a license verdict line with the verdict licensed.