在 Flashduty On-call
您可通过以下两种方式获取集成推送地址,任选其一即可。集成类型都选择 Google Cloud Monitoring(GCP 云监控),不是 Personalized Service Health。
使用专属集成
- 进入 Flashduty 控制台,选择 协作空间,打开一个协作空间
- 选择 配置 → 集成数据 → 专属集成,点击 新增一个集成
- 选择 GCP 云监控,点击 保存
- 打开生成的集成卡片,复制 推送地址,形如
https://api.flashcat.cloud/event/push/alert/google-cm?integration_key=<集成密钥>
使用共享集成
- 进入 Flashduty 控制台,选择 集成中心 → 告警事件
- 选择 GCP 云监控,填写集成名称
- 配置默认路由并选择协作空间;创建后可在 路由 中增加更多规则
- 点击 保存,复制生成的 推送地址
在 Google Cloud 中配置
步骤 1:准备
- 在要接收事件的项目中启用 Service Health API,并确认项目已启用结算
- 配置告警的账号需要项目上的这些 IAM 角色:Logs Configuration Writer(
roles/logging.configWriter)、Monitoring AlertPolicy Editor(roles/monitoring.alertPolicyEditor)、Monitoring NotificationChannel Viewer(roles/monitoring.notificationChannelViewer)、Personalized Service Health Viewer(roles/servicehealth.viewer)
步骤 2:创建 Webhook 通知渠道
- 在 Google Cloud 控制台进入 Monitoring → Alerting,点击 Edit notification channels
- 在 Webhook 部分点击 Add New,
Endpoint URL填入上一步复制的推送地址,Display Name填Flashduty - 点击 Test Connection 后点击 Save
步骤 3:创建 Service Health 告警策略
任选一种方式。 方式一:在 Service Health 控制台创建- 进入 Service Health 控制台,点击右上角 Create Alert Policy
- 选择一个告警策略模板,通知渠道选择步骤 2 创建的 Webhook 渠道,点击 Create Policies
- 需要修改条件时,在模板菜单中选择 Customize alert policy
policy.json 示例,NOTIFICATION_CHANNEL 是步骤 2 中渠道的资源名(形如 projects/PROJECT_ID/notificationChannels/885798905074,可用 gcloud beta monitoring channels list 查询):
filter 是日志过滤条件,Google 文档给出的常用条件:
产品 ID 和位置名称的取值见 Google 文档 Google Cloud products and locations。在告警策略中设置 Policy severity level(API 字段
severity),Flashduty 据此确定告警等级。
步骤 4:测试
按 Google 文档,向 Cloud Logging 写入一条测试日志,等待几分钟,在 Monitoring → Incidents 和 Flashduty 中确认收到告警。重复测试需要间隔至少 5 分钟。LOG_NAME 使用 Service Health 的日志名,Google 文档中的格式为 projects/PROJECT_ID/logs/servicehealth.googleapis.com%2Factivity。
开启超时自动关闭
对于日志告警策略,Cloud Monitoring 只在事件打开时发送通知,不会在关闭时发送通知(API 中日志告警策略的通知时机固定为
OPENED),所以 Flashduty 收不到恢复通知。请在接收这些告警的协作空间中开启 超时自动关闭,超时计时起点选择 故障触发,建议时长 24 小时,再按你所关心事件的常见持续时间调整。
告警策略里的 autoClose 只决定 Cloud Monitoring 中事件何时关闭(日志告警默认 7 天),不会通知 Flashduty。
字段映射
排查问题
- 没有收到告警:在 Monitoring → Incidents 确认事件已打开;检查日志过滤条件是否匹配,通知有频率限制(Google 文档中的示例为每个策略每天每个项目 20 条告警)
- 告警没有恢复:日志告警不发送关闭通知,请开启协作空间的超时自动关闭
- Test Connection 失败:确认
Endpoint URL是完整的推送地址,包含integration_key - 同一个事件的多次更新产生多条告警:每个 Cloud Monitoring 事件(
incident_id)对应一条 Flashduty 告警;用new_event条件只对新事件发出通知,可以减少告警数量