Skip to content

Monitoring

Describes Hibernator chart 0.12.44

Hibernator exposes Prometheus metrics on both components:

  • API: /metrics on port 8080
  • Controller: /metrics on port 9090

The chart can create a ServiceMonitor, a Grafana dashboard ConfigMap and a PrometheusRule for you. All three are off by default. Every value named below is in the table in configuration.md, under metrics.

metrics.serviceMonitor.enabled: true creates ServiceMonitor resources for both components, so a Prometheus Operator installation discovers them automatically. If your Prometheus selects ServiceMonitors by label — kube-prometheus-stack does, on release — put that label in metrics.serviceMonitor.additionalLabels, or the resources are created and never scraped.

Scrape interval and timeout follow the Prometheus defaults unless you set metrics.serviceMonitor.interval and metrics.serviceMonitor.scrapeTimeout.

metrics.grafanaDashboard.enabled: true ships the dashboard as a ConfigMap labelled for the Grafana sidecar to pick up. It has panels for build info, uptime, storage sizes, scaling operations, notification queue drops, reconciliation, API request rates, notification delivery (queue depth, send rate, retry rate) and infrastructure health (external wake circuit breaker state, ConfigMap conflict retries).

The sidecar label defaults to grafana_dashboard: "1"; change it with metrics.grafanaDashboard.sidecarLabel and .sidecarLabelValue if your Grafana looks for something else. metrics.grafanaDashboard.folder only takes effect if your provisioning has foldersFromFilesStructure enabled.

metrics.prometheusRule.enabled: true creates a PrometheusRule covering the failure modes worth paging on. Each alert can be turned off individually, and the thresholds are configurable through metrics.prometheusRule.alerts.<name>.thresholdBytes, .thresholdPercent and .thresholdSeconds.

Alert Severity Default threshold What it means
HibernatorScalingFailures critical any failure in 10m Hibernate or wake operations are failing
HibernatorConfigMapSizeCritical critical > 900KB State ConfigMap is approaching the Kubernetes 1MB limit
HibernatorParameterStoreSizeCritical critical > 3900B SSM parameter is approaching the 4KB limit
HibernatorConfigMapSizeWarning warning > 800KB State ConfigMap is growing large
HibernatorParameterStoreSizeWarning warning > 3500B SSM parameter is growing large
HibernatorPasskeyStoreSizeWarning warning > 600KB The passkey ConfigMap is growing large; passkeys are never evicted
HibernatorApiErrorRate warning > 5% 5xx The API is returning errors
HibernatorReconciliationStalled warning no completed pass in 2h, or no controller The controller’s reconcile loop has stopped, or the controller has no pod
HibernatorLicenseUnverified warning a day after the last licensed check The license check cannot tell; Hibernator still acts until the stop date in the description
HibernatorNotActing warning 5m Hibernator has stopped scaling: its license verdict does not let it act
HibernatorLicenseExpiring warning within 30 days The license expires soon

The three license alerts read the controller’s license metrics; none is critical, because a license issue stops scaling and never takes a workload down. license.md describes the metrics, events and notifications behind them, and what to do about each reason.

HibernatorPasskeyStoreSizeWarning fires below the generic ConfigMap warning on purpose. Passkeys are never evicted to make room, so the alert has to arrive while there is still space to act on it; at 600KB the store is holding several hundred users’ passkeys, against a ceiling near 500. passkeys.md says what to do about it. The alert is inert on installs that leave config.auth.webauthn.enabled at false, because the ConfigMap is never created.

HibernatorReconciliationStalled watches hibernator_reconciliation_last_success_timestamp_seconds, which the controller moves at the end of every reconciliation pass it completes. A pass completes when it acts, and also when it skips because a stand-down or the license check has stopped Hibernator: an inert controller is a live loop skipping on purpose, so neither raises the alert. hibernator_reconciliation_total{status="skipped"} counts those passes. The alert fires in two cases, each after its for::

  • the controller has completed no pass for longer than thresholdSeconds (7200 by default): the loop has stopped;
  • the series is missing: no controller pod is running, or Prometheus scrapes none.

Both cases match on the namespace label of the release, so when one Prometheus scrapes two installs, the healthy one cannot hide the one that stopped. The ServiceMonitor sets that label. If you scrape the controller another way, give it a namespace target label with the release namespace, or the alert fires for as long as the rule exists.

As with the ServiceMonitor, a label-selecting Prometheus needs metrics.prometheusRule.additionalLabels.

Three instruments measure what a wake waited for. All three are the controller’s alone — the API pod does not expose them — so scrape the controller for these, whatever else you scrape. None of the three has an alerting rule: they are diagnostic, for a question an operator asks after a wake, not paging signals.

A histogram with one observation per workload that arrived during a wake, measured from the pass that first saw that workload. It has no labels.

Arrived means the wake saw it short of ready at least once and then saw it ready. Every workload that does is in it: a wake a controller restart resumed contributes its workloads measured from the original wake’s start, downtime included, and a wake a hibernation superseded contributes the workloads that had already come ready. Leaving those out would tilt the distribution toward the happy path, which is the blindness the histogram exists to remove — the audit history it replaces stops at the settle deadline and records nothing beyond it, so it can show you a wake that wrote four workloads off at ten minutes but never that all four were ready at nineteen.

# Where the settle tail really sits.
histogram_quantile(0.9, sum by (le) (rate(hibernator_workload_settle_seconds_bucket[7d])))
# How many workloads settle after ten minutes (600 s).
sum(rate(hibernator_workload_settle_seconds_bucket{le="+Inf"}[7d]))
- sum(rate(hibernator_workload_settle_seconds_bucket{le="600"}[7d]))

Buckets run from 30 seconds to 40 minutes, densest between 4 and 12 minutes.

A workload that is already ready when the wake’s first pass looks at it is not an observation. It did not arrive during this wake — it was up before the wake started — and the histogram has no label to separate its near-zero reading from a fast cold start, so recording it would pull every quantile down on exactly the wakes that waited for nothing: a re-wake, one an external wake session triggered over a partly awake environment, a workload the cluster never scaled down. A wake where nothing had to start therefore contributes no observations at all, which is the honest answer rather than a pile of zeroes.

It also has no label to separate a recovered workload from a merely slow one. A workload hibernator stopped waiting for and then watched come up is one observation here like any other, so a rise in the tail can be either — the counter below is what tells them apart. That is deliberate rather than a gap: the histogram measures how long workloads took, one question, and a label splitting it by a verdict that was withdrawn would make every unlabelled query over it mean something subtly different. Know about it when you read the tail.

A counter of the workloads a wake stopped waiting for, labelled reason with one of nine codes. It counts only what a wake ended still holding: a workload written off at its deadline and observed ready before the wake settled has its verdict lifted, and is one observation of the histogram above plus one of the recoveries below, rather than one of these.

reason What hibernator saw What a rise means
image_pull Pods still failing to pull when the cluster’s budget for producing a started container ran out Your registry, or the path to it
unschedulable The cluster’s fifteen minutes ran out on a pod the scheduler had marked Unschedulable No node could take the pod: capacity, node selectors, taints, or the node autoscaler
pending The cluster’s fifteen minutes ran out on a pod with no node that the scheduler had not marked Unschedulable The scheduler: down, backlogged, or not the one the pod names
init_running The cluster’s fifteen minutes ran out on a pod whose init containers had not finished The init containers — a migration, a dependency they wait on
still_pulling The cluster’s fifteen minutes ran out on a pod whose container the kubelet was still creating, with no pull failing Slow pulls — node disk, image size. Not image_pull, which is a pull that keeps failing
deadline A budget spent with nothing terminal to point at: the application did not finish starting within the budget it declares, the kubelet refused to create a container, or the cluster’s fifteen minutes ran out on pods in none of the four states above Capacity — CPU contention on start-up; or, for a refused container, the manifest’s Secret and ConfigMap references
crash_loop A container that restarted twice before it was ready The application rejecting its own environment
unreadable A workload no readiness pass could read before its ceiling ran out Hibernator’s blind spot: RBAC, or an API server under load
invalid_image An image reference that is malformed, or absent under pull policy Never The manifest is wrong
# One alert per audience, rather than one alert for "a wake lost something".
sum by (reason) (increase(hibernator_workload_write_offs_total[1h]))

The codes are for querying, and they are not the words anyone reads. The completion dialog, the Teams card, the email and the audit entry each carry the full sentence the controller composed for that workload, naming the budget it was judged against — “did not finish starting within its declared 20 minutes”. A code groups those; it never replaces one.

A counter of the workloads a wake stopped waiting for and then observed ready anyway, labelled reason with the code of the write-off the lift cleared. It carries the same nine values as the counter above and the same label name, so the share that came back needs no label_replace:

# The share of write-offs that came back, by reason. A high ratio on `deadline` or
# `crash_loop` means your workloads are declaring less time than they need.
sum by (reason) (increase(hibernator_workload_recoveries_total[7d]))
/ (sum by (reason) (increase(hibernator_workload_recoveries_total[7d]))
+ (sum by (reason) (increase(hibernator_workload_write_offs_total[7d]))
or sum by (reason) (increase(hibernator_workload_recoveries_total[7d])) * 0))

Neither counter is pre-seeded — a series exists only once that reason has occurred — and the + matches one-to-one on reason, so the or … * 0 is what keeps a reason that has only ever recovered from dropping out of the result entirely. That is the 100 %-recovery case, which is exactly the one worth seeing. The other end is left alone: a reason that has only ever been written off is absent from the result rather than shown as 0, and absent reads as zero recoveries here. Add the mirror or … * 0 to the numerator if you would rather see the row.

The two partition, and both are taken at the end of the wake. A wake that ends still holding a write-off counts a casualty and no recovery, even where that workload’s verdict lifted and came back earlier in the same wake: the question each answers is what the wake ended with. So the two sum to at most one per workload the wake ever wrote off — never two, and neither where the wake ended while that workload was between verdicts, which it reports as still starting — and the ratio above is well formed.

A rise here is not an incident on its own — these are wakes that worked. What it measures is how often hibernator’s report was nearly wrong: crash_loop recovering routinely means applications are dying once on a cold cluster and coming back, and deadline recovering routinely means a startupProbe declaring less than the application needs. Both are worth a look at the application, not at hibernator.

metrics:
serviceMonitor:
enabled: true
additionalLabels:
release: kube-prometheus-stack
grafanaDashboard:
enabled: true
prometheusRule:
enabled: true
additionalLabels:
release: kube-prometheus-stack