Qeasy Cloud
Get Started

Bot Alert Notification Strategy in Practice: Pushing Integration Errors to DingTalk in Real Time

· 卢剑航· Integration Solutions· 7 views· 4 min read

What This Strategy Solves

In a supply chain integration running more than a dozen sync links, what really keeps us up at night is never a one-step delay — it is a strategy failing silently. Inventory stops pushing, orders get stuck in the middle, and by the time the business team complains, half a day is gone. We use the Qeasy data integration platform for this: alongside all the sync strategies, we add a dedicated "Bot Alert Notification" strategy that periodically queries error details from every relevant strategy and pushes anomalies to a DingTalk group, so ops and business users see them on their phones immediately.

Data Flow and Field Mapping

The data flow has three stages: strategy error details (source API) → assembly in the Qeasy middle layer → DingTalk robot Webhook (target API).

The source side calls Qeasy's own StrategyErrorDetail query API. Typical inputs:

FieldMeaningTypical Value
recentSecondsLookback window in seconds600
idsStrategy ID listMultiple strategy IDs separated by commas
statusError status code3 (means "error")

The output is filled by autoFillResponse. Core fields include strategy_name, strategy_id, number, response_at, and problem.

The middle layer rearranges these fields into the structure expected by a DingTalk message card, then calls the target DingTalkRobotDetail execution API. Typical mapping:

Source FieldTarget FieldNote
strategy_namenameName of the failing strategy, shown directly
lessee.namelessee_nameTenant name, to identify which account
numbernumberDocument number, for the business team to trace in the source system
response_atresponse_atWhen the error occurred
problemproblemError summary, the key line in the message body

The target side is only a DingTalk robot; it does not write to any business database, so the effect of this strategy is EXECUTE rather than QUERY.

How to Configure It on Qeasy

A few points are worth expanding when configuring this strategy on Qeasy:

  • Platform choice: Pick the custom WebAPI connector on both sides. The source points to StrategyErrorDetail and the target to DingTalkRobotDetail. Do not reuse a business-system connector here, because the source is querying Qeasy's own metadata API.
  • Request body: On the source, set buildModel=false and use the three fixed parameters directly. On the target, use template variables like {{strategy_name}} and {{number}} to pass through the fields returned from the previous step.
  • access_token: Place the DingTalk robot's access_token in the target platform's authentication configuration, not inline in the strategy script — it makes future rotation much easier.
  • Response handling: The target side does not depend on response data, so disable strict checks on autoFillResponse to avoid pointless errors.
  • Alert throttling: Add a rule in the Qeasy transformation component so the same strategy_id + same problem only fires once within 10 minutes. This is the key point revisited in the pitfalls section.

Implementation Steps

We recommend launching this strategy in three phases:

  1. Seed phase: Start with a single strategy ID and verify end-to-end that Qeasy returns errors and that a structured card shows up in the DingTalk group. Use a non-critical sync link (such as a logging sync) as the "seed strategy" so the first run does not trigger an alert storm.
  2. Full rollout: Once it works, change ids to the full list of strategy IDs that need monitoring. Use the "dual-track full + incremental" pattern common among Qeasy customers: the ids list acts as a full-coverage fallback, while status filtering lets new strategies join simply by adding an ID.
  3. Scheduling: Use */10 7-23 * * * on the source (every 10 minutes during business hours) and 2-59/10 7-23 * * * on the target, offset by 2 minutes so the two ends never contend for resources at the same second. Outside business hours, either disable the job or widen recentSeconds.

Pitfalls We Have Seen

  • Pitfall: alert storm. When a downstream API wobbles, 20 strategies error out within 10 minutes and the DingTalk group gets flooded. The safe approach is the "same strategy + same problem, once per 10 minutes" throttling rule mentioned above.
  • Pitfall: hard-coded ids missing new strategies. When a new strategy is launched, nobody remembers to add its ID into the alert strategy, so its errors never reach anyone. The safe approach is to externalize ids into a configuration table in Qeasy and make new strategies a dependency that auto-registers.
  • Pitfall: non-error statuses being pushed. The status field defaults to 3 when omitted, but in some cases "skipped schedule" (status 6) does not need an alert. Always pass status=3 explicitly rather than relying on the default.
  • Pitfall: exposed access_token. Writing the token directly into a strategy value means every strategy needs editing when the robot rotates. The safe approach is to centralize it in the connector's authentication configuration.
  • Pitfall: lookback window too short. Setting recentSeconds to 60 leaves a gap between two schedules and causes missed alerts. Keep it at 600–900 seconds, offset by half the schedule interval.

Suitable and Unsuitable Scenarios

Suitable: multi-strategy supply chain / finance integrations where a small-to-medium team needs 7×12-hour alert coverage, and the alert receiver is an IM that supports Webhooks (DingTalk, WeCom, Lark). Unsuitable: financial scenarios with sub-second alert latency requirements (IM push itself jitters), or environments with very large alert volumes that need tiered routing into PagerDuty / phone calls — those should integrate with a professional alerting platform instead of an IM robot.

Original content. Please credit the source when reposting: https://www.qeasy.cloud/insights/solutions/strat-jushuitan-kingdee-cloud-6235-na98d8460-54cb401a

Comments