
Error Handling
Part of Workflow error handling
Separating retryable failures from invalid data
Decide whether a failed workflow record needs a bounded retry, data correction or reconciliation of an uncertain write.
Ask what would have to change for the failed action to succeed. If time or dependency recovery could change the result, consider a bounded retry. If the record breaks a known rule, hold it for correction. If a write may already have succeeded, establish its outcome before repeating it.
Classify the observed failure
| Condition | Question | Initial route |
|---|---|---|
| Required value is absent or unsupported | Can an authorised source lookup supply the right value? | Hold for lookup or correction; do not invent a default. |
| Documented throttling or temporary unavailability | When may another attempt be made, and is it still useful? | Retry within the deadline if the action is safe to repeat. |
| Authentication or permission rejection | Is the connection or its access wrong? | Stop repeating the request and assign the connection owner. |
| Timeout after a create request was sent | Did the destination create the object? | Mark the outcome unknown and reconcile it. |
Use the selected service’s documented error response. A throttling response commonly signals a temporary dependency fault, while many invalid-request responses will not improve when the same payload is resent. A status code alone does not settle whether an earlier write took effect.
Retryable Failures vs Invalid Data: Key Differences
- Retryable Failure
- Temporary issue (e.g., throttling, timeout, transient dependency fault)
- Invalid Data
- Structural or business rule violation (e.g., missing required field, unsupported value)
Separate record errors from dependency errors
Validate the values required for the outbound operation. A missing JSON property differs from one explicitly set to null; the destination contract determines whether either is allowed. Structural validation can identify missing or mistyped fields. A business rule must decide whether a value, such as a delivery method, is appropriate for the record.
If an event lacks a required value, use an authorised source lookup where one exists. If no trustworthy value is available, record the failed rule and route the item for correction. Do not relabel an exhausted temporary failure as invalid data: the owner needs its actual diagnosis.
Put limits around a retry
For a likely temporary fault, set a per-call timeout, an attempt limit and an overall deadline. Follow a usable provider retry instruction. Check connector and SDK settings so a workflow attempt does not conceal several outbound calls. Preserve the action’s identity across attempts, and send work that exhausts its budget to an owned exception route.
Repeating a write requires its own safety check. For example, a hypothetical fulfilment API might create a dispatch request but lose the response. Look up the request using an agreed stable reference or use the destination’s documented safe retry mechanism if its conditions apply. If neither establishes the result, hold the record for investigation.
Check the classification rule
Before release, use safe cases for invalid input, a temporary outage, rejected access and a lost response after a write. Record the expected route, owner and destination state, then compare them with observed results. Include a corrected record so the route permits recovery after the original problem is resolved.


![How to choose the right idempotency key for an event: Keep the same idempotency key for retries of one action; use a new key for a new action.; Write: one [effect] for each [identity], e.g. one standard invoice per approved order.; Check stability, uniqueness, scope and lifetime; include tenant ID for tenant-scoped IDs. Choosing an idempotency key for an event](/covers/choosing-an-idempotency-key-for-an-event-640.webp?v=39704930)
