> ## Documentation Index
> Fetch the complete documentation index at: https://docs.flashduty.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Azure Service Health 告警集成

> 通过 Service Health 告警规则和 Action group 的 Webhook 接收 Azure 服务事件，使用 Flashduty 的 Azure Monitor 集成，事件解决后自动恢复告警。

Azure 把服务问题、计划内维护、运行状况公告和安全公告作为 Service Health 通知写入活动日志，Service Health 告警规则可以通过 Action group 的 Webhook 把它们发出去。Flashduty 的 [Azure Monitor 集成](/zh/on-call/integration/alert-integration/alert-sources/azure-monitor) 能解析 Service Health 通知，同一个事件（相同订阅和 Tracking ID）合并为一条告警，因此不需要单独的 Azure Service Health 集成：在 Flashduty 创建 Azure Monitor 集成，把它的推送地址填进 Action group 即可。

<div className="hide">
  ## 在 Flashduty On-call

  ***

  您可通过以下两种方式获取集成推送地址，任选其一即可。**集成类型都选择 Azure Monitor**，不是 Azure Service Health。

  ### 使用专属集成

  1. 进入 Flashduty 控制台，选择 **协作空间**，打开一个协作空间
  2. 选择 **配置** → **集成数据** → **专属集成**，点击 **新增一个集成**
  3. 选择 **Azure Monitor**，点击 **保存**
  4. 打开生成的集成卡片，复制 **推送地址**

  ### 使用共享集成

  1. 进入 Flashduty 控制台，选择 **集成中心 → 告警事件**
  2. 选择 **Azure Monitor**，填写集成名称
  3. 配置默认路由并选择协作空间；创建后可在 **路由** 中增加更多规则
  4. 点击 **保存**，复制生成的 **推送地址**
</div>

## 在 Azure 中配置

***

### 步骤 1：创建带 Webhook 的 Action group

1. 在 Azure 门户进入 **Monitor**，打开 **Alerts** → **Action groups**，创建 Action group
2. 在 **Actions** 中添加 **Webhook** 类型的动作，`URI` 填入上一步复制的推送地址
3. 启用通用告警架构（common alert schema）后保存。未启用时 Flashduty 也能解析 Service Health 的旧版负载，区别见下文“标题和描述”

<Note>
  Azure 要求：Service Health 告警使用的 Action group，其区域必须设置为 **Global**；Service Health 告警只支持全球区域的公有云。
</Note>

### 步骤 2：创建 Service Health 告警规则

1. 在 Azure 门户选择 **Service Health**，在 **Service Issues** 面板点击 **Create service health alert**
2. 在 **Scope** 中按订阅选择范围
3. 在 **Condition** 中选择 **Services**、**Regions** 和 **Event types**（服务问题、计划内维护、运行状况公告、安全公告）。Azure 建议同时选择全部服务和区域：只有影响你所用区域的事件才会触发告警，不会因未使用的服务产生告警
4. 在 **Details** 中选择资源组并填写告警规则名称
5. 点击 **Advanced options**，在 **Actions** 标签页选择步骤 1 创建的 Action group，然后创建

## 字段映射

***

| Service Health 字段 | Flashduty |
| :- | :- |
| 订阅 ID + `properties.trackingId` | Alert Key。同一订阅下同一个 Azure 事件的后续通知合并为一条告警 |
| `properties.stage` | 阶段为 `Resolved`、`RCA`、`Complete`、`Completed` 或 `Canceled` 时恢复告警，其他阶段（例如 `Active`、`Planned`）触发或更新告警 |
| `essentials.severity`（通用架构） | `Sev0` → Critical；`Sev1`、`Sev2` → Warning；`Sev3`、`Sev4` → Info。未启用通用架构时统一为 Info |
| `properties.service`、`region`、`incidentType` | 标签 `service`、`region`、`incident_type` |
| `properties.impactedServices` | 标签 `impacted_services`、`impacted_regions` |
| `properties.stage`、`trackingId` | 标签 `stage`、`tracking_id` |
| 订阅 ID + `trackingId` | 标签 `service_health_url`，指向该事件的 Service Health 页面 |

## 标题和描述

***

| | 启用通用告警架构 | 未启用（旧版 `Microsoft.Insights/activityLogs`） |
| :- | :- | :- |
| 告警标题 | 告警规则名称（订阅级规则没有资源，只显示规则名称） | 事件标题（`properties.title`，缺失时取 `defaultLanguageTitle`） |
| 告警描述 | 无 | Azure 通报正文（`properties.communication`，缺失时取 `defaultLanguageContent`，去除 HTML 标签） |
| Alert Key 与恢复 | 相同 | 相同 |

通用架构下，事件标题和正文保留在 `alertContext` 标签的 JSON 中。需要在通知里直接看到事件标题时，使用旧版负载（不启用通用告警架构）。

## 恢复与去重

***

* Azure 在事件进展时向同一个 Tracking ID 发送多条通知，Flashduty 按 Alert Key 合并；事件进入解决阶段时恢复告警
* 缺少 `properties.trackingId` 的旧版通知会被拒绝（HTTP 400）；通用架构下缺少 Tracking ID 的通知每条单独成为一条告警，不会自动恢复
* 计划内维护在 `Complete` 或 `Canceled` 阶段恢复；在维护开始前，告警处于触发状态

## 排查问题

***

* **Action group 没有触发**：确认 Action group 的区域是 **Global**，且告警规则的 **Actions** 标签页已选择该 Action group
* **告警没有恢复**：确认告警规则仍然启用，并等待 Azure 发出事件解决阶段的通知
* **Azure 显示 Webhook 调用失败**：确认 `URI` 是完整的推送地址，包含 `integration_key`
* **看到多条告警**：不同的 Azure 事件有不同的 Tracking ID，各自对应一条告警；同一事件的进展通知会合并

更多字段含义请参阅 Azure 文档 [Create Service Health alerts](https://learn.microsoft.com/azure/service-health/alerts-activity-log-service-notifications-portal)、[Activity log alert webhook schema](https://learn.microsoft.com/azure/azure-monitor/alerts/activity-log-alerts-webhook) 和 [Common alert schema](https://learn.microsoft.com/azure/azure-monitor/alerts/alerts-common-schema)。
