Bot Alert Notification Strategy in Practice: Pushing Integration Errors to DingTalk in Real Time
What This Strategy Solves
In a supply chain integration running more than a dozen sync links, what really keeps us up at night is never a one-step delay — it is a strategy failing silently. Inventory stops pushing, orders get stuck in the middle, and by the time the business team complains, half a day is gone. We use the Qeasy data integration platform for this: alongside all the sync strategies, we add a dedicated "Bot Alert Notification" strategy that periodically queries error details from every relevant strategy and pushes anomalies to a DingTalk group, so ops and business users see them on their phones immediately.
Data Flow and Field Mapping
The data flow has three stages: strategy error details (source API) → assembly in the Qeasy middle layer → DingTalk robot Webhook (target API).
The source side calls Qeasy's own StrategyErrorDetail query API. Typical inputs:
| Field | Meaning | Typical Value |
|---|---|---|
| recentSeconds | Lookback window in seconds | 600 |
| ids | Strategy ID list | Multiple strategy IDs separated by commas |
| status | Error status code | 3 (means "error") |
The output is filled by autoFillResponse. Core fields include strategy_name, strategy_id, number, response_at, and problem.
The middle layer rearranges these fields into the structure expected by a DingTalk message card, then calls the target DingTalkRobotDetail execution API. Typical mapping:
| Source Field | Target Field | Note |
|---|---|---|
| strategy_name | name | Name of the failing strategy, shown directly |
| lessee.name | lessee_name | Tenant name, to identify which account |
| number | number | Document number, for the business team to trace in the source system |
| response_at | response_at | When the error occurred |
| problem | problem | Error summary, the key line in the message body |
The target side is only a DingTalk robot; it does not write to any business database, so the effect of this strategy is EXECUTE rather than QUERY.
How to Configure It on Qeasy
A few points are worth expanding when configuring this strategy on Qeasy:
- Platform choice: Pick the custom WebAPI connector on both sides. The source points to
StrategyErrorDetailand the target toDingTalkRobotDetail. Do not reuse a business-system connector here, because the source is querying Qeasy's own metadata API. - Request body: On the source, set
buildModel=falseand use the three fixed parameters directly. On the target, use template variables like{{strategy_name}}and{{number}}to pass through the fields returned from the previous step. - access_token: Place the DingTalk robot's access_token in the target platform's authentication configuration, not inline in the strategy script — it makes future rotation much easier.
- Response handling: The target side does not depend on response data, so disable strict checks on
autoFillResponseto avoid pointless errors. - Alert throttling: Add a rule in the Qeasy transformation component so the same
strategy_id+ sameproblemonly fires once within 10 minutes. This is the key point revisited in the pitfalls section.
Implementation Steps
We recommend launching this strategy in three phases:
- Seed phase: Start with a single strategy ID and verify end-to-end that Qeasy returns errors and that a structured card shows up in the DingTalk group. Use a non-critical sync link (such as a logging sync) as the "seed strategy" so the first run does not trigger an alert storm.
- Full rollout: Once it works, change
idsto the full list of strategy IDs that need monitoring. Use the "dual-track full + incremental" pattern common among Qeasy customers: theidslist acts as a full-coverage fallback, while status filtering lets new strategies join simply by adding an ID. - Scheduling: Use
*/10 7-23 * * *on the source (every 10 minutes during business hours) and2-59/10 7-23 * * *on the target, offset by 2 minutes so the two ends never contend for resources at the same second. Outside business hours, either disable the job or widenrecentSeconds.
Pitfalls We Have Seen
- Pitfall: alert storm. When a downstream API wobbles, 20 strategies error out within 10 minutes and the DingTalk group gets flooded. The safe approach is the "same strategy + same problem, once per 10 minutes" throttling rule mentioned above.
- Pitfall: hard-coded
idsmissing new strategies. When a new strategy is launched, nobody remembers to add its ID into the alert strategy, so its errors never reach anyone. The safe approach is to externalizeidsinto a configuration table in Qeasy and make new strategies a dependency that auto-registers. - Pitfall: non-error statuses being pushed. The
statusfield defaults to 3 when omitted, but in some cases "skipped schedule" (status 6) does not need an alert. Always passstatus=3explicitly rather than relying on the default. - Pitfall: exposed access_token. Writing the token directly into a strategy value means every strategy needs editing when the robot rotates. The safe approach is to centralize it in the connector's authentication configuration.
- Pitfall: lookback window too short. Setting
recentSecondsto 60 leaves a gap between two schedules and causes missed alerts. Keep it at 600–900 seconds, offset by half the schedule interval.
Suitable and Unsuitable Scenarios
Suitable: multi-strategy supply chain / finance integrations where a small-to-medium team needs 7×12-hour alert coverage, and the alert receiver is an IM that supports Webhooks (DingTalk, WeCom, Lark). Unsuitable: financial scenarios with sub-second alert latency requirements (IM push itself jitters), or environments with very large alert volumes that need tiered routing into PagerDuty / phone calls — those should integrate with a professional alerting platform instead of an IM robot.