
Error Handling
Part of Stateful workflow orchestration
Resuming a workflow after a service interruption
Use the last confirmed outcome and destination records to resume an interrupted workflow without blindly repeating work.
Resume an interrupted workflow from its last confirmed outcome. A durable runtime may restore progress after a worker stops, but an external request sent just before the interruption may already have succeeded. Resolve that unknown outcome before repeating a write.
Identify the interruption
Find the instance ID and inspect its execution history: current stage, completed task, pending event or timer, and error. Distinguish a stopped worker from a failed external task or an execution that has ended unsuccessfully. A log entry saying a request was sent establishes an attempt, not its destination outcome.
If the instance is waiting normally, restoring a worker may let the runtime continue. A failed task may follow a configured retry or recovery route. An unsuccessful execution may need an explicit restart or redrive if its platform supports one.
Reconcile the external action
Classify each outbound action as confirmed success, confirmed failure without an applied effect, or unknown. Treat an error or timeout as unknown when it does not establish whether the destination committed the write. For confirmed success, retain the destination object's ID.
For an unknown write, search by a stable business reference or use a documented idempotency mechanism whose scope and retention cover the retry. If neither is available, stop automated retries and assign investigation.
Suppose a workflow submitted a purchase order but lost the response. The purchasing application may already hold the order. Find it by the agreed reference and record its ID before continuing. Submitting again merely because the workflow still says submitting risks creating a duplicate.
Resume the appropriate execution
Check whether the platform will replay history, retry a task or redrive an execution. In Temporal, a worker replays recorded workflow history to rebuild state, and completed activity results in that history are returned during replay. An activity attempt can still be retried, so its external effect needs protection.
AWS Step Functions can redrive eligible unsuccessful Standard Workflow executions within its documented 14-day period and execution-history limit. It preserves earlier successful steps, but reruns the unsuccessful Task state. Redrive therefore still requires reconciliation of a write whose outcome is unknown. It uses the execution's existing state-machine definition, which matters if the definition has since changed.
Before continuation, check that the current business record still permits the next action. A request may have been cancelled, changed or superseded during the interruption.
Key Recovery Metrics and Limits
- AWS Step Functions Redrive Window
- 14 days
- Temporal Replay Capability
- Yes – based on recorded workflow history
- Idempotency Requirement
- Mandatory for external writes with unknown outcomes
Close the recovery
Compare the workflow's final state with the destination record. Record affected instance IDs, last confirmed effects, unresolved cases and actions taken.
Keep unresolved instances visible until an owner closes them. A useful recovery exercise covers interruption before a side effect, after an effect but before its response is recorded, and during a wait.



![How to choose the right idempotency key for an event: Keep the same idempotency key for retries of one action; use a new key for a new action.; Write: one [effect] for each [identity], e.g. one standard invoice per approved order.; Check stability, uniqueness, scope and lifetime; include tenant ID for tenant-scoped IDs. Choosing an idempotency key for an event](/covers/choosing-an-idempotency-key-for-an-event-640.webp?v=39704930)