Skip to main content
When a workflow instance ends or times out, Apache DolphinScheduler sends a notification to the alert instances in the selected alarm group. If you point an alert instance of the HTTP plugin at the Flashduty push URL, a failed workflow instance creates one Critical alert in Flashduty, and the alert recovers when the same instance is re-run and succeeds. This integration was verified against DolphinScheduler 3.4.3. You must fill in Body as shown in Configure the body, or DolphinScheduler sends no alert content.

In Flashduty On-call


You can get the integration push URL in either of two ways. Pick one.

Use a dedicated integration

  1. Go to the Flashduty console, select Channels, and open a channel
  2. Select Settings → Integration data → Dedicated integrations, then click Add an integration
  3. Select Dolphinscheduler, then click Save
  4. Open the new integration card and copy the push URL

Use a shared integration

  1. Go to the Flashduty console and select Integration Center → Alert events
  2. Select Dolphinscheduler and enter an integration name
  3. Configure the default route and pick a channel; you can add more rules under Routes after it is created
  4. Click Save and copy the generated push URL

In DolphinScheduler


1

Create an HTTP alarm instance

  1. Sign in to DolphinScheduler with an account that has permission, go to Security → Alarm Instance Manage, and click Create Alarm Instance
  2. For Select plugin choose Http, and enter an Alarm instance name
  3. Fill in the plugin parameters:
  1. Save. The DolphinScheduler server must be able to reach Flashduty on the public internet.
Configure the body: the HTTP plugin sends only what you put in Body, and replaces ${msg} inside a string value of the body with the alert content. Without ${msg}, Flashduty receives no alert content. The value of content must be "${msg}"; other keys are ignored. After replacement, content is a JSON string that holds an array of alert objects, and Flashduty parses it a second time.
2

Create an alarm group

  1. Go to Security → Alarm Group Manage and click Create Alarm Group
  2. Enter an Alert Group Name, select the alarm instance from the previous step under Alarm Plugin Instance, and save
3

Select the alarm group and notification strategy on the workflow

When you start a workflow, or set a schedule for it:
  1. Select the alarm group from the previous step for Alarm Group
  2. Set Notification Strategy to All (ALL). With Failure, only failures are sent and alerts in Flashduty never recover
4

Verify

In Alarm Instance Manage, open the alarm instance for editing, click Test Send in the dialog, and confirm Flashduty shows one Info alert titled DolphinScheduler test notification. Then make a task in a workflow fail (for example a Shell task that runs exit 1) and confirm a Critical alert titled DolphinScheduler workflow failed: <workflow instance name> appears.

Recovery and auto-close


DolphinScheduler sends a notification according to the state the workflow instance ends in: When you re-run a failed instance with Recovery Failed (START_FAILURE_TASK_PROCESS), the instance ID stays the same and runTimes increases by 1. If the re-run fails again, the same alert is updated. If it succeeds and the notification strategy is All, the alert recovers. Workflow and task timeout alerts are sent once and never recover. If nobody re-runs a failed workflow, no recovery is sent either. In the channel that receives this integration, turn on auto-close. We suggest 12 hours; adjust to how quickly your team handles failed workflows. Sub-workflow states are not notified on their own.

Alert types


The content array of one request can hold several alert objects. Each object becomes one Flashduty alert.
  • Workflow instance failure or success: the title is DolphinScheduler workflow failed: <workflow instance name> or DolphinScheduler workflow succeeded: <workflow instance name>
  • Workflow timeout: the title is DolphinScheduler workflow timeout: <workflow instance name>, Warning severity
  • Task timeout: the title is DolphinScheduler task timeout: <task name>, Warning severity
  • Test send: the test message of an alarm instance is fixed. Flashduty recognizes it and creates a separate Info alert; every press is a new alert and never merges with or closes a real alert. Close it by hand
If the alarm group also receives alerts that carry no workflow information (for example a service-down alert), Flashduty rejects them with an invalid-parameter error. To keep failed sends out of DolphinScheduler, create a separate alarm group for this integration.

Alert Key


Changes to the workflow instance name, state, run count, or times do not change the Alert Key. A failure and the later success of the same instance share one Alert Key, so the success recovers the failure. An alert object without projectCode or workflowInstanceId makes the whole request be rejected.

Status and severity


Labels


Troubleshooting


  • Flashduty receives no events: confirm the workflow was started with an alarm group and a notification strategy other than None, that the alarm group contains the alarm instance, and that the push URL is complete and includes integration_key. Test Send in the alarm instance’s create or edit dialog checks the network path
  • Flashduty returns an invalid-parameter error: usually Body is not {"content":"${msg}"}, or an alert without workflow information (for example service-down) was received
  • An alert never closes: the failed instance was not re-run successfully, or the notification strategy is not All. Turn on auto-close for the channel
  • No success notifications: the notification strategy is Failure