
Error Handling
Part of Workflow testing and deployment
Rolling back a workflow without replaying completed actions
Redirect new workflow starts, classify affected outcomes and recover unfinished items without repeating confirmed external actions.
Redirect new work to an approved version, then reconcile work affected by the faulty release before retrying it. Restoring an older definition does not undo invoices, messages or records already created by a newer one.
Stop new exposure
Identify the faulty version, when it began receiving work and how new executions start. Pause that route or direct new starts to the approved version using the platform's documented control. Preserve affected execution IDs and history for investigation.
In AWS Step Functions, an alias can route new executions back to an earlier published version when callers start through that alias. A caller using an unqualified state-machine ARN instead starts the latest revision.
Where a platform provides a documented process for returning to an earlier design, follow that process and verify the actual start route after the change.
Classify work from the release window
List the items that entered while the faulty version was active. For each intended external action, record its source reference, available destination or provider ID, and one current state:
| State | Next decision |
|---|---|
| Confirmed complete | Retain the result; do not repeat the action. |
| Confirmed no effect | Recheck current eligibility before controlled processing. |
| Accepted, outcome pending | Check the destination's later result before retrying. |
| Outcome unknown | Look up the effect by a stable reference or use a documented safe retry method; otherwise hold for investigation. |
A failed workflow status cannot prove an external create did nothing. The destination may have committed the write before its response was lost.
Handle existing executions separately
Changing the version for new starts may leave existing executions on their original definition. Step Functions redrive applies only to eligible unsuccessful Standard Workflow executions and uses the original definition even after an alias changes.
It normally preserves successful steps and reruns the unsuccessful Task. The documentation includes a section on redrive behaviour of individual states, so inspect the failed state and any external effects before redrive.
A version restore does not compensate for a wrong message or record. The business owner must decide whether the destination needs a correction, cancellation or other compensating action. Repeating the old workflow is not automatically that correction.
Workflow rollback: What can and cannot be undone
- Can be recoveredUnfinished tasks, pending outcomes, and recheckable eligibility after failure.
- Cannot be undoneInvoices, messages, or records already created by the faulty version; these require manual correction.
Recover only unfinished obligations
Match eligible source items from the release window to confirmed destination results. For each gap, check current source state: an item may have changed or been cancelled. Process only still-eligible, uncompleted work through an authorised recovery route that retains the action's stable identity. Keep confirmed results closed.
Record the restored version, affected window, classifications, recovery decisions and remaining owners. The rollback review is complete when each eligible item has a confirmed outcome or an assigned exception, and a fresh authorised input follows the intended restored route.
Post-rollback verification checklist
- ✓ Redirected new executions to approved versionVerified via alias or routing control in platform (e.g., AWS Step Functions).
- ✓ Classified all affected work from release windowApplied state-based decisions per documented classification table.
- ✓ Redriven only eligible failed executionsConfirmed compatibility with redrive rules in AWS Step Functions.
- ✓ Documented recovery decisions and ownersRetained records of restored version, affected window, and assigned exceptions.



![How to choose the right idempotency key for an event: Keep the same idempotency key for retries of one action; use a new key for a new action.; Write: one [effect] for each [identity], e.g. one standard invoice per approved order.; Check stability, uniqueness, scope and lifetime; include tenant ID for tenant-scoped IDs. Choosing an idempotency key for an event](/covers/choosing-an-idempotency-key-for-an-event-640.webp?v=39704930)