Routing failed records to an exception queue: Send invalid records for correction; queue temporary failures when their budget expires.; Assign an owner and review deadline; storage alone does not recover the work.; Service Bus DLQ ignores TTL and never auto-cleans; SQS redrive resets retention.
Image: Workflow Automation Guide

Error Handling

Part of Workflow error handling

Routing failed records to an exception queue

Keep failed workflow records diagnosable and owned, then check retention, destination state and replay safety before recovery.

Route a record when normal processing cannot finish within its retry budget or a person must correct it. Keep enough context to find the record, understand its last confirmed outcome and choose the next action.

Assign an owner and a review deadline; storage alone does not recover the work.

From failed record to resolved exception

  1. Detect failureNormal processing cannot finish within its retry budget, or a person must correct it.
  2. ClassifyKnown invalid: correction. Likely temporary: attempt or time budget expired. Uncertain outcome: reconciliation, marked outcome unknown.
  3. RouteSend to the exception queue with enough context to find the record and understand its last confirmed outcome.
  4. RecordCapture source reference, intended action, last confirmed stage, destination result ID and safe failure details.
  5. OwnAssign an owner and review deadline; define who may inspect, correct, replay and close entries.
  6. ReviewCheck exception store retention and message age before the work expires or becomes too late to correct.
  7. ResolveCorrect the source, connection or workflow rule through its owner and confirm the action is still eligible.
  8. ReplayReplay through the agreed action identity, verify the destination result and record the resolution.

Decide what enters

Send known invalid records for correction without repeatedly calling the destination. Send likely temporary failures when their attempt or time budget expires. Send writes with uncertain outcomes for reconciliation, marked outcome unknown.

A broker’s automatic dead-letter rule does not decide whether a failure needs correction or reconciliation. Azure Service Bus provides a dead-letter queue for messages that cannot be delivered or processed; Amazon SQS moves messages under a configured receive-count policy. Record why the intended business action failed.

Keep an operable exception record

Include the stable source reference, intended action, last confirmed stage and any destination result ID. Record the failure category, safe error code, attempt history, owner, status and next permitted action.

Keep credentials and unnecessary personal information out of the entry. Where a source reference is enough, avoid copying the full payload; retain the context needed to identify the failed version if the source later changes.

Assign an owner and review deadline, and define who may inspect, correct, replay and close entries. Do not acknowledge the intake message until its required outcome or a durable exception hand-off is confirmed.

In a Service Bus peek-lock flow, use the parent entity’s dead-letter operation to hand off a message. Retrieve and handle the DLQ message, then complete it; Service Bus also supports receive-and-delete delivery.

Operable exception record checklist

  • Stable source reference
  • Intended action
  • Last confirmed stage
  • Destination result ID (if any)
  • Failure category
  • Safe error code
  • Attempt history
  • Owner
  • Status
  • Next permitted action
  • No credentials or unnecessary personal information
  • Context needed to identify the failed version if the source later changes

Set retention and review timing

Check how long the chosen exception store keeps work and how it measures message age. Azure Service Bus does not observe time-to-live in its dead-letter queue and does not automatically clean it up; messages remain until retrieved and completed.

Amazon SQS redrive resets the retention period, and redriven messages are considered new messages. Review work before it expires or becomes too late to correct.

Dead-letter retention and redrive behaviour

Azure Service Bus: dead-letter trigger
Messages that cannot be delivered or processed
Azure Service Bus: time-to-live in DLQ
Not observed in the dead-letter queue
Azure Service Bus: automatic cleanup
No automatic cleanup; messages remain until retrieved and completed
Amazon SQS: dead-letter trigger
Configured receive-count policy
Amazon SQS: redrive retention
Retention period resets; redriven messages are considered new messages
Amazon SQS: default redrive destination
Moves messages from a DLQ to the source queue by default
Amazon SQS: destination choice
Can choose another destination queue of the same type
Amazon SQS: move rate
Can set the move rate
Amazon SQS: API action
`StartMessageMoveTask` starts an asynchronous redrive

Replay after resolving the cause

Inspect the authoritative source record and check whether the destination already contains the intended effect. Correct the source, connection or workflow rule through its owner, then confirm the action is still eligible.

Replay through a route that preserves the agreed action identity, verify the destination result and record the resolution. For Service Bus, inspect the DLQ message and use an application route to correct and resubmit it; complete the DLQ message after handling.

Amazon SQS redrive moves messages from a DLQ to the source queue by default. You can choose another destination queue of the same type and set the move rate; the API action StartMessageMoveTask starts an asynchronous redrive.

If correction requires a changed payload or selection of individual records, use a separately designed recovery route. Do not act on a changed or cancelled source record merely because its old message remains in the queue.

After a shared outage, try a small set first, then pace the remaining replay to the destination’s capacity. Keep each exception visible as an individual record.

More from Error Handling