Monitoring
Describes Hibernator chart 0.12.44
Hibernator exposes Prometheus metrics on both components:
- API:
/metricson port 8080 - Controller:
/metricson port 9090
The chart can create a ServiceMonitor, a Grafana dashboard ConfigMap and a
PrometheusRule for you. All three are off by default. Every value named below is in the
table in configuration.md, under metrics.
- ServiceMonitor
- Grafana dashboard
- Alerting rules
- Wake settle, write-offs and recoveries
- Turning all three on
ServiceMonitor
Section titled “ServiceMonitor”metrics.serviceMonitor.enabled: true creates ServiceMonitor resources for both
components, so a Prometheus Operator installation discovers them automatically. If your
Prometheus selects ServiceMonitors by label — kube-prometheus-stack does, on
release — put that label in metrics.serviceMonitor.additionalLabels, or the
resources are created and never scraped.
Scrape interval and timeout follow the Prometheus defaults unless you set
metrics.serviceMonitor.interval and metrics.serviceMonitor.scrapeTimeout.
Grafana dashboard
Section titled “Grafana dashboard”metrics.grafanaDashboard.enabled: true ships the dashboard as a ConfigMap labelled
for the Grafana sidecar to pick up. It has
panels for build info, uptime, storage sizes, scaling operations, notification queue
drops, reconciliation, API request rates, notification delivery (queue depth, send
rate, retry rate) and infrastructure health (external wake circuit breaker state,
ConfigMap conflict retries).
The sidecar label defaults to grafana_dashboard: "1"; change it with
metrics.grafanaDashboard.sidecarLabel and .sidecarLabelValue if your Grafana looks
for something else. metrics.grafanaDashboard.folder only takes effect if your
provisioning has foldersFromFilesStructure enabled.
Alerting rules
Section titled “Alerting rules”metrics.prometheusRule.enabled: true creates a PrometheusRule covering the failure
modes worth paging on. Each alert can be turned off individually, and the thresholds
are configurable through metrics.prometheusRule.alerts.<name>.thresholdBytes,
.thresholdPercent and .thresholdSeconds.
| Alert | Severity | Default threshold | What it means |
|---|---|---|---|
HibernatorScalingFailures |
critical | any failure in 10m | Hibernate or wake operations are failing |
HibernatorConfigMapSizeCritical |
critical | > 900KB | State ConfigMap is approaching the Kubernetes 1MB limit |
HibernatorParameterStoreSizeCritical |
critical | > 3900B | SSM parameter is approaching the 4KB limit |
HibernatorConfigMapSizeWarning |
warning | > 800KB | State ConfigMap is growing large |
HibernatorParameterStoreSizeWarning |
warning | > 3500B | SSM parameter is growing large |
HibernatorPasskeyStoreSizeWarning |
warning | > 600KB | The passkey ConfigMap is growing large; passkeys are never evicted |
HibernatorApiErrorRate |
warning | > 5% 5xx | The API is returning errors |
HibernatorReconciliationStalled |
warning | no completed pass in 2h, or no controller | The controller’s reconcile loop has stopped, or the controller has no pod |
HibernatorLicenseUnverified |
warning | a day after the last licensed check | The license check cannot tell; Hibernator still acts until the stop date in the description |
HibernatorNotActing |
warning | 5m | Hibernator has stopped scaling: its license verdict does not let it act |
HibernatorLicenseExpiring |
warning | within 30 days | The license expires soon |
The three license alerts read the controller’s license metrics; none is critical, because a license issue stops scaling and never takes a workload down. license.md describes the metrics, events and notifications behind them, and what to do about each reason.
HibernatorPasskeyStoreSizeWarning fires below the generic ConfigMap warning on
purpose. Passkeys are never evicted to make room, so the alert has to arrive while
there is still space to act on it; at 600KB the store is holding several hundred
users’ passkeys, against a ceiling near 500. passkeys.md
says what to do about it. The alert is inert on installs that leave
config.auth.webauthn.enabled at false, because the ConfigMap is never created.
HibernatorReconciliationStalled watches hibernator_reconciliation_last_success_timestamp_seconds,
which the controller moves at the end of every reconciliation pass it completes. A pass
completes when it acts, and also when it skips because a stand-down or the license check
has stopped Hibernator: an inert controller is a live loop skipping on purpose, so
neither raises the alert. hibernator_reconciliation_total{status="skipped"} counts
those passes. The alert fires in two cases, each after its for::
- the controller has completed no pass for longer than
thresholdSeconds(7200 by default): the loop has stopped; - the series is missing: no controller pod is running, or Prometheus scrapes none.
Both cases match on the namespace label of the release, so when one Prometheus scrapes
two installs, the healthy one cannot hide the one that stopped. The ServiceMonitor sets
that label. If you scrape the controller another way, give it a namespace target label
with the release namespace, or the alert fires for as long as the rule exists.
As with the ServiceMonitor, a label-selecting Prometheus needs
metrics.prometheusRule.additionalLabels.
Wake settle, write-offs and recoveries
Section titled “Wake settle, write-offs and recoveries”Three instruments measure what a wake waited for. All three are the controller’s alone — the API pod does not expose them — so scrape the controller for these, whatever else you scrape. None of the three has an alerting rule: they are diagnostic, for a question an operator asks after a wake, not paging signals.
hibernator_workload_settle_seconds
Section titled “hibernator_workload_settle_seconds”A histogram with one observation per workload that arrived during a wake, measured from the pass that first saw that workload. It has no labels.
Arrived means the wake saw it short of ready at least once and then saw it ready. Every workload that does is in it: a wake a controller restart resumed contributes its workloads measured from the original wake’s start, downtime included, and a wake a hibernation superseded contributes the workloads that had already come ready. Leaving those out would tilt the distribution toward the happy path, which is the blindness the histogram exists to remove — the audit history it replaces stops at the settle deadline and records nothing beyond it, so it can show you a wake that wrote four workloads off at ten minutes but never that all four were ready at nineteen.
# Where the settle tail really sits.histogram_quantile(0.9, sum by (le) (rate(hibernator_workload_settle_seconds_bucket[7d])))
# How many workloads settle after ten minutes (600 s).sum(rate(hibernator_workload_settle_seconds_bucket{le="+Inf"}[7d])) - sum(rate(hibernator_workload_settle_seconds_bucket{le="600"}[7d]))Buckets run from 30 seconds to 40 minutes, densest between 4 and 12 minutes.
A workload that is already ready when the wake’s first pass looks at it is not an observation. It did not arrive during this wake — it was up before the wake started — and the histogram has no label to separate its near-zero reading from a fast cold start, so recording it would pull every quantile down on exactly the wakes that waited for nothing: a re-wake, one an external wake session triggered over a partly awake environment, a workload the cluster never scaled down. A wake where nothing had to start therefore contributes no observations at all, which is the honest answer rather than a pile of zeroes.
It also has no label to separate a recovered workload from a merely slow one. A workload hibernator stopped waiting for and then watched come up is one observation here like any other, so a rise in the tail can be either — the counter below is what tells them apart. That is deliberate rather than a gap: the histogram measures how long workloads took, one question, and a label splitting it by a verdict that was withdrawn would make every unlabelled query over it mean something subtly different. Know about it when you read the tail.
hibernator_workload_write_offs_total
Section titled “hibernator_workload_write_offs_total”A counter of the workloads a wake stopped waiting for, labelled reason with one of nine
codes. It counts only what a wake ended still holding: a workload written off at its
deadline and observed ready before the wake settled has its verdict lifted, and is one
observation of the histogram above plus one of the recoveries below, rather than one of
these.
reason |
What hibernator saw | What a rise means |
|---|---|---|
image_pull |
Pods still failing to pull when the cluster’s budget for producing a started container ran out | Your registry, or the path to it |
unschedulable |
The cluster’s fifteen minutes ran out on a pod the scheduler had marked Unschedulable | No node could take the pod: capacity, node selectors, taints, or the node autoscaler |
pending |
The cluster’s fifteen minutes ran out on a pod with no node that the scheduler had not marked Unschedulable | The scheduler: down, backlogged, or not the one the pod names |
init_running |
The cluster’s fifteen minutes ran out on a pod whose init containers had not finished | The init containers — a migration, a dependency they wait on |
still_pulling |
The cluster’s fifteen minutes ran out on a pod whose container the kubelet was still creating, with no pull failing | Slow pulls — node disk, image size. Not image_pull, which is a pull that keeps failing |
deadline |
A budget spent with nothing terminal to point at: the application did not finish starting within the budget it declares, the kubelet refused to create a container, or the cluster’s fifteen minutes ran out on pods in none of the four states above | Capacity — CPU contention on start-up; or, for a refused container, the manifest’s Secret and ConfigMap references |
crash_loop |
A container that restarted twice before it was ready | The application rejecting its own environment |
unreadable |
A workload no readiness pass could read before its ceiling ran out | Hibernator’s blind spot: RBAC, or an API server under load |
invalid_image |
An image reference that is malformed, or absent under pull policy Never |
The manifest is wrong |
# One alert per audience, rather than one alert for "a wake lost something".sum by (reason) (increase(hibernator_workload_write_offs_total[1h]))The codes are for querying, and they are not the words anyone reads. The completion dialog, the Teams card, the email and the audit entry each carry the full sentence the controller composed for that workload, naming the budget it was judged against — “did not finish starting within its declared 20 minutes”. A code groups those; it never replaces one.
hibernator_workload_recoveries_total
Section titled “hibernator_workload_recoveries_total”A counter of the workloads a wake stopped waiting for and then observed ready anyway,
labelled reason with the code of the write-off the lift cleared. It carries the same nine
values as the counter above and the same label name, so the share that came back needs no
label_replace:
# The share of write-offs that came back, by reason. A high ratio on `deadline` or# `crash_loop` means your workloads are declaring less time than they need.sum by (reason) (increase(hibernator_workload_recoveries_total[7d])) / (sum by (reason) (increase(hibernator_workload_recoveries_total[7d])) + (sum by (reason) (increase(hibernator_workload_write_offs_total[7d])) or sum by (reason) (increase(hibernator_workload_recoveries_total[7d])) * 0))Neither counter is pre-seeded — a series exists only once that reason has occurred — and the
+ matches one-to-one on reason, so the or … * 0 is what keeps a reason that has only
ever recovered from dropping out of the result entirely. That is the 100 %-recovery case,
which is exactly the one worth seeing. The other end is left alone: a reason that has only
ever been written off is absent from the result rather than shown as 0, and absent reads as
zero recoveries here. Add the mirror or … * 0 to the numerator if you would rather see the
row.
The two partition, and both are taken at the end of the wake. A wake that ends still holding a write-off counts a casualty and no recovery, even where that workload’s verdict lifted and came back earlier in the same wake: the question each answers is what the wake ended with. So the two sum to at most one per workload the wake ever wrote off — never two, and neither where the wake ended while that workload was between verdicts, which it reports as still starting — and the ratio above is well formed.
A rise here is not an incident on its own — these are wakes that worked. What it measures is
how often hibernator’s report was nearly wrong: crash_loop recovering routinely means
applications are dying once on a cold cluster and coming back, and deadline recovering
routinely means a startupProbe declaring less than the application needs. Both are worth a
look at the application, not at hibernator.
Turning all three on
Section titled “Turning all three on”metrics: serviceMonitor: enabled: true additionalLabels: release: kube-prometheus-stack grafanaDashboard: enabled: true prometheusRule: enabled: true additionalLabels: release: kube-prometheus-stack