Resuming workflows after interruptions: Check instance ID and execution history to identify the interruption type.; Reconcile external actions using business references or idempotency mechanisms.; Verify current state and record outcomes before resuming with ATO-compliant logging.
Image: Workflow Automation Guide

Error Handling

Part of Stateful workflow orchestration

Resuming a workflow after a service interruption

Use the last confirmed outcome and destination records to resume an interrupted workflow without blindly repeating work.

Resume an interrupted workflow from its last confirmed outcome. A durable runtime may restore progress after a worker stops, but an external request sent just before the interruption may already have succeeded. Resolve that unknown outcome before repeating a write.

Identify the interruption

Find the instance ID and inspect its execution history: current stage, completed task, pending event or timer, and error. Distinguish a stopped worker from a failed external task or an execution that has ended unsuccessfully. A log entry saying a request was sent establishes an attempt, not its destination outcome.

If the instance is waiting normally, restoring a worker may let the runtime continue. A failed task may follow a configured retry or recovery route. An unsuccessful execution may need an explicit restart or redrive if its platform supports one.

Reconcile the external action

Classify each outbound action as confirmed success, confirmed failure without an applied effect, or unknown. Treat an error or timeout as unknown when it does not establish whether the destination committed the write. For confirmed success, retain the destination object's ID.

For an unknown write, search by a stable business reference or use a documented idempotency mechanism whose scope and retention cover the retry. If neither is available, stop automated retries and assign investigation.

Suppose a workflow submitted a purchase order but lost the response. The purchasing application may already hold the order. Find it by the agreed reference and record its ID before continuing. Submitting again merely because the workflow still says submitting risks creating a duplicate.

Resume the appropriate execution

Check whether the platform will replay history, retry a task or redrive an execution. In Temporal, a worker replays recorded workflow history to rebuild state, and completed activity results in that history are returned during replay. An activity attempt can still be retried, so its external effect needs protection.

AWS Step Functions can redrive eligible unsuccessful Standard Workflow executions within its documented 14-day period and execution-history limit. It preserves earlier successful steps, but reruns the unsuccessful Task state. Redrive therefore still requires reconciliation of a write whose outcome is unknown. It uses the execution's existing state-machine definition, which matters if the definition has since changed.

Before continuation, check that the current business record still permits the next action. A request may have been cancelled, changed or superseded during the interruption.

Key Recovery Metrics and Limits

AWS Step Functions Redrive Window
14 days
Temporal Replay Capability
Yes – based on recorded workflow history
Idempotency Requirement
Mandatory for external writes with unknown outcomes

Close the recovery

Compare the workflow's final state with the destination record. Record affected instance IDs, last confirmed effects, unresolved cases and actions taken.

Keep unresolved instances visible until an owner closes them. A useful recovery exercise covers interruption before a side effect, after an effect but before its response is recorded, and during a wait.

More from Error Handling