Skip to content

Notifications

Describes Hibernator chart 0.12.44

Hibernator sends a card to Microsoft Teams, or an email, when something happens in the cluster. Examples are a wake, a hibernation, a failure, a license change and a stand-down. Notifications are off by default.

The controller finds an event and gives it to the API, and the API pod sends the card or the email. A destination is one email recipient or one Teams webhook.

  • config.notifications.events is the default for every destination. A destination’s own events map wins for each key that it sets (Choose the events for a destination).
  • The API reads these values when it starts. When a helm upgrade changes one of them, the chart restarts the API pod.
  • config.clusterName is required. Every card and email names the cluster with it, and the chart refuses notifications.enabled: true without it.
  1. Add the notification values to your values file. This example sends to one Teams channel and one email address:

    config:
    clusterName: "production-eu"
    notifications:
    enabled: true
    channels:
    teams:
    enabled: true
    webhooks:
    - name: "Operations"
    webhookURL: "https://example.com/your-teams-workflow-url"
    events:
    high: true
    medium: true
    email:
    enabled: true
    recipients:
    - email: "platform-team@example.com"
  2. Apply the values file. The chart restarts the API pod, and the new pod reads the new values:

    Terminal window
    helm upgrade hibernator oci://registry.gitlab.com/cirriton/hibernator/charts/hibernator \
    -n hibernator -f my-values.yaml
  3. Wait until the new API pod is ready. Then check the API log. The line Initializing notification service shows that the service started:

    Terminal window
    kubectl logs -n hibernator deployment/hibernator-app -c api | grep -i notification

A recipient or a webhook without an events map gets the values of config.notifications.events and the defaults below.

Each event has a key and a priority: high, medium or low. For each key of a destination, the chart takes the first of these that it finds:

  1. The key in the destination’s events map, for example scheduledWake: true.
  2. The key in config.notifications.events.
  3. The key of its priority (high, medium or low) in the destination’s events map.
  4. The key of its priority in config.notifications.events.
  5. The default of the key.

A key in config.notifications.events therefore wins over a priority key in the destination’s map. In this example, the first recipient gets the scheduled wakes, and the second does not:

config:
notifications:
events:
scheduledWake: true
channels:
email:
recipients:
- email: "platform-team@example.com"
events:
low: false
- email: "on-call@example.com"
events:
scheduledWake: false

confirmationWindowSeconds is only a global value. A destination’s events map does not take it.

Events Default for an email recipient Default for a Teams webhook
high on on
medium off on
recoveryCatchup, licenseUnverified, licenseNotActing, licenseRestored, licenseExpiring, standDown (medium) on on
low off off

medium: false therefore also turns off the license, stand-down and catch-up events. To keep one of them, set its key to true as well.

A wake that lost workloads goes further. Its card is red, and it also goes to every destination that takes at least one high-priority event. This applies also when the destination does not take the wake’s own key, for example scheduledWake.

Four warnings go to every destination, and no key turns them off. They report a ConfigMap that is too large, a Parameter Store parameter that is too large, an invalid schedule and a wake barrier timeout.

Key Priority Card When
manualScaleUp high Wake Complete a wake that was not scheduled completes: a manual wake, an external wake or a handback
manualScaleDown high Manual Scale Down a hibernation that the schedule did not start completes, for example after Trigger Manual Hibernation
hibernationFailed high Hibernation Failed a workload does not scale down, or a database does not stop
wakeFailed high Wake Failed a workload does not scale up, or a database does not start
configReloadFailed high Configuration Reload Failed the controller cannot load a changed configuration
externalWakeStarted medium External Wake Started, External Wake Renewed a dependent cluster starts or renews an external wake session on this cluster
externalWakeExpired medium External Wake Expired an external wake session ends
unexpectedReplicaChange medium Unexpected Replica Change something other than Hibernator scales a running workload to 0, or up from 0, and the change stays for the confirmation window
stateMismatch medium State Mismatch Detected something scales up a hibernated workload without a wake, and the change stays for the confirmation window
resourceAdded medium Resource Added Hibernator starts to manage a workload
resourceRemoved medium Resource Removed Hibernator stops managing a workload
scheduleOverride medium Schedule Override someone changes the schedule in the UI
recoveryCatchup medium Recovery Catch-up the controller does a hibernation or a wake that it missed while it was down; this card replaces the scheduled one
licenseUnverified medium License Check Failing the license check cannot tell, and grace begins; again 3 days before the stop date
licenseNotActing medium Stopped Scaling the license check stops Hibernator scaling
licenseRestored medium License Check Succeeded a license check succeeds again
licenseExpiring medium License Expiring the license expires in 30 days, and again in 3
standDown medium Stood Down, Stand-down Ended a stand-down begins or ends
scheduledHibernation low Scheduled Hibernation the schedule hibernates the cluster
scheduledWake low Wake Complete the schedule wakes the cluster
databaseStopped low Database Stopped Hibernator stops managed databases, one stop delay after a hibernation

License and Stop Hibernator in an emergency tell more about their cards. External wake tells about its sessions.

The email channel sends through the same mail server as the sign-in codes: config.smtp and secrets.smtp (Authentication). The emails show your brand’s logo (Branding).

  • The channel has no mail server settings of its own. The chart refuses smtpHost, smtpPort, from, smtpUsername and smtpPassword under channels.email.
  • Set the server and the sender in config.smtp. Set the credentials in secrets.smtp, or in your own Secret if secrets.existingSecret names one.
  • The chart refuses channels.email.enabled: true while config.smtp.enabled is false.
  • The chart refuses an address in recipients that is not a valid email address.

Each webhook posts to one Teams channel through a Teams workflow.

  • name is required, and Settings → Configuration shows it.
  • webhookURL is required. It is the URL of the workflow. The URL is signed, and anyone who has it can post to the channel. The chart puts it in the configuration ConfigMap, not in a Secret.
  • channels.teams.proxyURL sends the cards through a proxy. When it is empty, the cards use config.proxyURL.

Teams cards show no logo.

  • Queue. Notifications wait in a queue in the API until they are sent. queue.bufferSize is its size. When the queue is full, the API drops the notification and logs Notification queue full, event dropped.
  • Deduplication. deduplication.windowSeconds is the time during which Hibernator does not send a repeated condition alert again. A report of one event, such as a completed wake, is always sent. Hibernator uses 300 when the value is 0. Troubleshooting tells which cards the window holds back.
  • Confirmation window. events.confirmationWindowSeconds applies to unexpectedReplicaChange and stateMismatch only. Hibernator sends the card only if the workload is still in the unexpected state when the window ends. With the value 0, Hibernator sends the card at once.
  • Retries. When a send fails, the API tries again, up to retry.maxAttempts tries for each channel. Before try n, it waits (n − 1) × retry.backoffSeconds seconds. All tries of one notification on one channel must end within retry.timeoutSeconds.

All values are under config.notifications.

Value Default What it does
enabled false Turns notifications on. Needs config.clusterName
events.confirmationWindowSeconds 300 Seconds that a workload must stay in an unexpected state before unexpectedReplicaChange or stateMismatch is sent. 0 sends at once
events.<key>, events.high, events.medium, events.low not set The default for each destination that does not set the key itself
queue.bufferSize 20 Notifications that can wait in the queue. At least 1
deduplication.windowSeconds 300 Seconds during which a repeated condition alert is not sent again. Not negative
retry.maxAttempts 3 Tries for each channel. At least 1
retry.backoffSeconds 2 Wait before try n: (n − 1) × this value, in seconds. At least 1
retry.timeoutSeconds 30 Time for all tries of one notification on one channel. At least 1
channels.email.enabled false Turns the email channel on. Needs at least one recipient and config.smtp.enabled: true
channels.email.recipients[].email The address. Required
channels.email.recipients[].events not set The events this address gets. A key that it does not set comes from events
channels.teams.enabled false Turns the Teams channel on. Needs at least one webhook
channels.teams.proxyURL not set Proxy for the Teams cards. Empty means config.proxyURL
channels.teams.webhooks[].name The name of the webhook. Required
channels.teams.webhooks[].webhookURL The URL of the Teams workflow. Required
channels.teams.webhooks[].events not set The events this webhook gets. A key that it does not set comes from events