Scrabble tiles on wood form 'FAIL', symbolizing defeat and reflection.
Photo by Markus Winkler on Pexels

Error Handling

Workflow error handling

Classify workflow failures, limit safe retries, preserve unresolved records and confirm the destination outcome.

Handle a workflow failure by deciding whether the record needs correction, can be retried safely, or has an unknown destination outcome. Give each route a limit and an owner. A failed workflow run does not always mean the destination did nothing.

Consider a hypothetical paid order sent to a fulfilment system. An unsupported delivery method needs correction. A temporary service failure may clear. A timeout after a create request may mean the fulfilment request already exists.

Define the result and the failure state

For each outbound action, identify what confirms success: a destination record ID, a completed state or another result the destination exposes. If an API only accepts work for later processing, keep acceptance separate from the final outcome.

Record the source reference, attempted action and last confirmed result. Use the source application for current business facts; the error record should explain the workflow's progress and who owns the next decision.

Route each failure

What is knownNext action
Input breaks a known ruleHold it for an authorised lookup or correction.
A dependency appears temporarily unavailableRetry within an attempt and time budget, if repeating the action is safe.
A write may have reached the destination, but its response was lostCheck the destination or use its documented safe retry mechanism before another create attempt.
Access or configuration is brokenStop repeated requests and assign the connection owner.

Use the selected API's error contract, not an HTTP status alone, to classify the failure. Throttling and some server errors are common retry candidates; repeating an unchanged invalid request usually cannot fix it.

Recognise transient conditions

Transient faults can arise from momentary network connectivity loss, a temporarily unavailable service or a timeout while a service is busy. They can occur across platforms and operating environments, so a workflow that communicates with remote services needs a recovery path rather than treating every interruption as permanent.

In cloud environments, shared resources may be throttled to protect them. A service can refuse new connections when its load reaches a limit or maximum throughput, allowing it to process existing requests and maintain performance for other users. A delayed retry may succeed once the temporary condition clears; repeated immediate requests can add pressure instead of helping.

Bound retries and preserve unresolved work

Check whether the connector or client already retries. Set a timeout for each call, a maximum attempt count and a business deadline. Honour a usable provider retry instruction.

Otherwise, delayed retries with random variation can help avoid simultaneous attempts by many workers. Count outbound calls across retry layers.

When a record cannot finish within that budget, put it on an owned exception route. A broker dead-letter queue and a purpose-built review store have different retention and replay behaviour; confirm the rules of the chosen service. Keep the record reference, failure category, attempts, last known outcome and next action. Avoid copying credentials or unnecessary personal information into the exception.

An owner may need to correct, reject or reconcile the record. Before replay, check current source eligibility and any destination effect that may already exist. Close the exception when the required result is confirmed or the record is deliberately rejected.

Alert and verify recovery

Notify the responsible owner when a record needs action. Before release, check representative invalid input, temporary outage, exhausted retries and lost-response cases in a safe environment. Compare the expected destination and exception states with what is observed. In operation, reconcile eligible source records against confirmed destination results and open exceptions.

In this guide

  1. Separating retryable failures from invalid dataDecide whether a failed workflow record needs a bounded retry, data correction or reconciliation of an uncertain write.
  2. Routing failed records to an exception queueKeep failed workflow records diagnosable and owned, then check retention, destination state and replay safety before recovery.
  3. Alerting an owner without creating an alert stormMake workflow failure alerts actionable, group related records and escalate urgent work without flooding owners.

More from Error Handling