warning or critical, and recovers automatically when the check returns to passing.
This integration works with both Consul Community Edition and Consul Enterprise. It needs only a watch definition on one Consul agent, with no extra scripts or plugins.
In Flashduty On-call
You can obtain an integration push URL in either of the following ways.
Use a dedicated integration
Choose this method when you do not need to route alerts to different channels.Expand
Expand
- In the Flashduty console, select Channel and open a channel
- Select Configuration → Integrations → Private integration, then click Add an integration
- Select HashiCorp Consul, then click Save
- Open the generated integration card and copy the Push URL
Use a shared integration
Choose this method when you need to route alerts to different channels based on the payload.Expand
Expand
- In the Flashduty console, select Integration Center → Alert Events
- Select HashiCorp Consul and enter an integration name
- Configure the default route and select a channel; after creation, add more rules under Route if needed
- Click Save and copy the generated Push URL
How it works
Whenever any watched health check changes (its status or its output), a Consul watch of type
checks POSTs the current state of all watched checks to the push URL as one JSON array. Flashduty turns each check in the array into one event:
- A
warningorcriticalcheck raises an alert, or merges into the alert it already has - A
passingcheck recovers its alert; if there is no such alert, nothing happens - The
_node_maintenanceand_service_maintenance:<service ID>checks created by maintenance mode (consul maint) are planned changes and never create alerts
warning or critical raise alerts right away.
Prerequisites
- Network: the Consul agent that runs the watch must be able to reach the Flashduty push URL (for example
https://api.flashcat.cloud). - Permissions: you need to be able to edit that agent’s configuration directory (for example
/etc/consul.d) and runconsul reload. - ACL: when ACLs are enabled, the token the watch uses needs read access to all nodes and services, as described below.
In Consul
1
Choose the agent that runs the watch
The watch queries the health checks of the whole datacenter, so configure it on one agent: a server, or a dedicated client. Configuring the same watch on several agents does not create duplicate alerts (events for the same check merge), but it multiplies the event count of every alert.
2
Add the watch definition
In that agent’s configuration directory, create a file named
flashduty-watch.json with the content below, and replace path with the full push URL of your Flashduty integration (including the integration_key parameter):-
Do not set the
stateparameter. With"state": "critical", a check drops out of the array when it recovers, Flashduty never receivespassing, and the alert cannot recover automatically. -
To watch only some checks, narrow the watch with the
filterparameter, for example"filter": "ServiceName == \"web\" or CheckID == \"serfHealth\"", or watch the checks of one service with"service": "web"(this leaves out node-level checks such asserfHealth). -
When ACLs are enabled, add
"token": "<ACL token>"to the watch. The token’s policy needs at least:
3
Reload the configuration
/consul/config in the container, then run:http watch handler failed with output, Flashduty rejected the request; the log line carries the HTTP status and the reason.4
Verify an alert and its recovery
Consul has no test notification button. Register a TTL check to verify the setup instead. A TTL check starts as Set the check to Deregister the check when you are done:When ACLs are enabled, add
critical, so it raises an alert right away:passing, and the alert recovers:-H "X-Consul-Token: <ACL token>" to each command.Alert Key
Flashduty builds the Alert Key from
Partition, Node, and CheckID.
Nodeis the name of the node the check runs on.CheckIDis the ID of the check. Consul requiresCheckIDto be unique on a node; a service check’s default ID looks likeservice:<service ID>, and the node liveness check isserfHealth.Partitionis the admin partition in Consul Enterprise and is empty in Community Edition.
Status and severity
When a check moves between
warning and critical, Flashduty keeps one alert per severity: warning turning critical creates a new Critical alert, and passing recovers both.
Labels
The alert title is
<check name> on <node name>, and the alert description is the check’s Output and Notes.
FAQ
A check or node was removed and its alert did not recover. What should I do?
A check or node was removed and its alert did not recover. What should I do?
When a
warning or critical check is deregistered, or its node leaves the cluster, it no longer appears in the array the watch sends, so Flashduty never receives a recovery. Close the alert by hand. In environments where nodes are removed often, turn on the auto-resolve timeout in the channel.Does Consul retry a failed delivery?
Does Consul retry a failed delivery?
No. Consul’s HTTP handler does not retry a failed request. The next time any watched check changes, the watch sends the current state of every check again, which fills in a missed trigger or recovery.
Why does the event count of an alert keep growing?
Why does the event count of an alert keep growing?
Every change to any watched check, including a change in output, makes the watch send all checks, and each failing check merges one more event into its alert. Use the
filter or service parameter to watch only the checks your on-call team acts on.The agent log shows request body exceeds 1 MiB limit. What should I do?
The agent log shows request body exceeds 1 MiB limit. What should I do?
One delivery must not exceed 1 MiB. With many checks, split them across several watches with the
filter or service parameter; all watches can use the same push URL.Do I need consul-alerts, as in the PagerDuty integration?
Do I need consul-alerts, as in the PagerDuty integration?
No. consul-alerts is a third-party daemon; this integration uses Consul’s own watch mechanism and needs no extra component.