
Error Handling
Part of Workflow error handling
Routing failed records to an exception queue
Keep failed workflow records diagnosable and owned, then check retention, destination state and replay safety before recovery.
Route a record when normal processing cannot finish within its retry budget or a person must correct it. Keep enough context to find the record, understand its last confirmed outcome and choose the next action.
Assign an owner and a review deadline; storage alone does not recover the work.
From failed record to resolved exception
- Detect failureNormal processing cannot finish within its retry budget, or a person must correct it.
- ClassifyKnown invalid: correction. Likely temporary: attempt or time budget expired. Uncertain outcome: reconciliation, marked outcome unknown.
- RouteSend to the exception queue with enough context to find the record and understand its last confirmed outcome.
- RecordCapture source reference, intended action, last confirmed stage, destination result ID and safe failure details.
- OwnAssign an owner and review deadline; define who may inspect, correct, replay and close entries.
- ReviewCheck exception store retention and message age before the work expires or becomes too late to correct.
- ResolveCorrect the source, connection or workflow rule through its owner and confirm the action is still eligible.
- ReplayReplay through the agreed action identity, verify the destination result and record the resolution.
Decide what enters
Send known invalid records for correction without repeatedly calling the destination. Send likely temporary failures when their attempt or time budget expires. Send writes with uncertain outcomes for reconciliation, marked outcome unknown.
A broker’s automatic dead-letter rule does not decide whether a failure needs correction or reconciliation. Azure Service Bus provides a dead-letter queue for messages that cannot be delivered or processed; Amazon SQS moves messages under a configured receive-count policy. Record why the intended business action failed.
Keep an operable exception record
Include the stable source reference, intended action, last confirmed stage and any destination result ID. Record the failure category, safe error code, attempt history, owner, status and next permitted action.
Keep credentials and unnecessary personal information out of the entry. Where a source reference is enough, avoid copying the full payload; retain the context needed to identify the failed version if the source later changes.
Assign an owner and review deadline, and define who may inspect, correct, replay and close entries. Do not acknowledge the intake message until its required outcome or a durable exception hand-off is confirmed.
In a Service Bus peek-lock flow, use the parent entity’s dead-letter operation to hand off a message. Retrieve and handle the DLQ message, then complete it; Service Bus also supports receive-and-delete delivery.
Operable exception record checklist
- Stable source reference
- Intended action
- Last confirmed stage
- Destination result ID (if any)
- Failure category
- Safe error code
- Attempt history
- Owner
- Status
- Next permitted action
- No credentials or unnecessary personal information
- Context needed to identify the failed version if the source later changes
Set retention and review timing
Check how long the chosen exception store keeps work and how it measures message age. Azure Service Bus does not observe time-to-live in its dead-letter queue and does not automatically clean it up; messages remain until retrieved and completed.
Amazon SQS redrive resets the retention period, and redriven messages are considered new messages. Review work before it expires or becomes too late to correct.
Dead-letter retention and redrive behaviour
- Azure Service Bus: dead-letter trigger
- Messages that cannot be delivered or processed
- Azure Service Bus: time-to-live in DLQ
- Not observed in the dead-letter queue
- Azure Service Bus: automatic cleanup
- No automatic cleanup; messages remain until retrieved and completed
- Amazon SQS: dead-letter trigger
- Configured receive-count policy
- Amazon SQS: redrive retention
- Retention period resets; redriven messages are considered new messages
- Amazon SQS: default redrive destination
- Moves messages from a DLQ to the source queue by default
- Amazon SQS: destination choice
- Can choose another destination queue of the same type
- Amazon SQS: move rate
- Can set the move rate
- Amazon SQS: API action
- `StartMessageMoveTask` starts an asynchronous redrive
Replay after resolving the cause
Inspect the authoritative source record and check whether the destination already contains the intended effect. Correct the source, connection or workflow rule through its owner, then confirm the action is still eligible.
Replay through a route that preserves the agreed action identity, verify the destination result and record the resolution. For Service Bus, inspect the DLQ message and use an application route to correct and resubmit it; complete the DLQ message after handling.
Amazon SQS redrive moves messages from a DLQ to the source queue by default. You can choose another destination queue of the same type and set the move rate; the API action StartMessageMoveTask starts an asynchronous redrive.
If correction requires a changed payload or selection of individual records, use a separately designed recovery route. Do not act on a changed or cancelled source record merely because its old message remains in the queue.
After a shared outage, try a small set first, then pace the remaining replay to the destination’s capacity. Keep each exception visible as an individual record.


![How to choose the right idempotency key for an event: Keep the same idempotency key for retries of one action; use a new key for a new action.; Write: one [effect] for each [identity], e.g. one standard invoice per approved order.; Check stability, uniqueness, scope and lifetime; include tenant ID for tenant-scoped IDs. Choosing an idempotency key for an event](/covers/choosing-an-idempotency-key-for-an-event-640.webp?v=39704930)
