
Error Handling
Workflow error handling
Classify workflow failures, limit safe retries, preserve unresolved records and confirm the destination outcome.
Handle a workflow failure by deciding whether the record needs correction, can be retried safely, or has an unknown destination outcome. Give each route a limit and an owner. A failed workflow run does not always mean the destination did nothing.
Consider a hypothetical paid order sent to a fulfilment system. An unsupported delivery method needs correction. A temporary service failure may clear. A timeout after a create request may mean the fulfilment request already exists.
Define the result and the failure state
For each outbound action, identify what confirms success: a destination record ID, a completed state or another result the destination exposes. If an API only accepts work for later processing, keep acceptance separate from the final outcome.
Record the source reference, attempted action and last confirmed result. Use the source application for current business facts; the error record should explain the workflow's progress and who owns the next decision.
Route each failure
| What is known | Next action |
|---|---|
| Input breaks a known rule | Hold it for an authorised lookup or correction. |
| A dependency appears temporarily unavailable | Retry within an attempt and time budget, if repeating the action is safe. |
| A write may have reached the destination, but its response was lost | Check the destination or use its documented safe retry mechanism before another create attempt. |
| Access or configuration is broken | Stop repeated requests and assign the connection owner. |
Use the selected API's error contract, not an HTTP status alone, to classify the failure. Throttling and some server errors are common retry candidates; repeating an unchanged invalid request usually cannot fix it.
Recognise transient conditions
Transient faults can arise from momentary network connectivity loss, a temporarily unavailable service or a timeout while a service is busy. They can occur across platforms and operating environments, so a workflow that communicates with remote services needs a recovery path rather than treating every interruption as permanent.
In cloud environments, shared resources may be throttled to protect them. A service can refuse new connections when its load reaches a limit or maximum throughput, allowing it to process existing requests and maintain performance for other users. A delayed retry may succeed once the temporary condition clears; repeated immediate requests can add pressure instead of helping.
Bound retries and preserve unresolved work
Check whether the connector or client already retries. Set a timeout for each call, a maximum attempt count and a business deadline. Honour a usable provider retry instruction.
Otherwise, delayed retries with random variation can help avoid simultaneous attempts by many workers. Count outbound calls across retry layers.
When a record cannot finish within that budget, put it on an owned exception route. A broker dead-letter queue and a purpose-built review store have different retention and replay behaviour; confirm the rules of the chosen service. Keep the record reference, failure category, attempts, last known outcome and next action. Avoid copying credentials or unnecessary personal information into the exception.
An owner may need to correct, reject or reconcile the record. Before replay, check current source eligibility and any destination effect that may already exist. Close the exception when the required result is confirmed or the record is deliberately rejected.
Alert and verify recovery
Notify the responsible owner when a record needs action. Before release, check representative invalid input, temporary outage, exhausted retries and lost-response cases in a safe environment. Compare the expected destination and exception states with what is observed. In operation, reconcile eligible source records against confirmed destination results and open exceptions.
In this guide
- Separating retryable failures from invalid dataDecide whether a failed workflow record needs a bounded retry, data correction or reconciliation of an uncertain write.
- Routing failed records to an exception queueKeep failed workflow records diagnosable and owned, then check retention, destination state and replay safety before recovery.
- Alerting an owner without creating an alert stormMake workflow failure alerts actionable, group related records and escalate urgent work without flooding owners.
![How to choose the right idempotency key for an event: Keep the same idempotency key for retries of one action; use a new key for a new action.; Write: one [effect] for each [identity], e.g. one standard invoice per approved order.; Check stability, uniqueness, scope and lifetime; include tenant ID for tenant-scoped IDs. Choosing an idempotency key for an event](/covers/choosing-an-idempotency-key-for-an-event-640.webp?v=39704930)


