
Error Handling
Part of Workflow monitoring and observability
Recording enough context to diagnose a failed workflow
Capture the IDs, attempted action, known outcome and safe error context needed to investigate a failed workflow.
A useful failure record tells an operator which business item was affected, which action was attempted, what outcome is known and where to check next. A stack trace without a source reference is hard to act on. Copying an entire payload into a log can expose data the investigation does not need.
Record the attempted action and its outcome
Identify the workflow and version, environment, source-item reference, run or instance ID, step, operation, attempt number and timestamp. Record a safe error category, provider request ID when available, last confirmed stage and destination result ID when known. Separate confirmed, rejected, accepted for later processing and unknown outbound outcomes where the operation permits those distinctions.
This illustrative event describes a lost response; it is not a provider payload:
{
"time": "2026-10-01T02:15:00Z",
"workflow": "order_to_fulfilment",
"workflow_version": "demo_v3",
"environment": "test",
"source_item_id": "order_demo_42",
"run_id": "run_demo_7",
"step": "create_fulfilment_request",
"attempt": 2,
"error_class": "response_timeout",
"destination_result_id": null,
"outcome": "unknown"
}
Here, null means no destination ID has been confirmed in this record. It does not establish that the destination created nothing. Use the agreed source reference or a documented destination lookup before another create attempt.
Outcome classifications for workflow actions
- Confirmed
- Action successfully completed and verified in destination.
- Rejected
- Action failed with explicit rejection; no further processing expected.
- Accepted for later processing
- Action accepted but not yet complete; may be delayed or retried.
- Unknown
- No confirmation of outcome; requires investigation.
Key identifiers to record in a failed workflow log
- Workflow name
- order_to_fulfilment
- Environment
- test
Preserve the path through systems
Carry a correlation reference across the source event, workflow and downstream call where supported. Keep the provider's request or message ID separately; it identifies that boundary rather than the business item. If tracing is available, include trace and span IDs in relevant logs.
OpenTelemetry's logging specification describes how those IDs correlate logs with traces from the same execution context. A trace shows the path of an attempt, not proof that an external write took effect.
For delayed or retried work, retain when the source event occurred, when each attempt started and when its result became known. Keep the last confirmed stage instead of replacing it with the latest error. Record the work's trigger and processing status, and use the outcome distinctions above to clarify what is known about the destination write.
Limit sensitive detail
Use safe identifiers and rule or error codes where they answer the diagnostic question. Keep passwords, tokens, connection strings and unnecessary personal information out of logs. Sanitise untrusted error text and restrict access to diagnostic records.
OWASP's logging guidance advises not logging too much or too little. It notes the operational and security uses of application logs.
Check what the workflow platform itself retains. Azure Logic Apps documentation includes guidance on hiding inputs, outputs, passwords and secrets in run history, and on securing run history. Check the documented controls and inspect the actual actions and logging route.
Balancing diagnostic detail with data security
- Pros of recording detailed contextEnables faster diagnosis, reduces manual investigation, supports audit trails.
- Cons of excessive loggingExposes sensitive data, increases risk of breaches, violates privacy regulations like GDPR and Australia’s Privacy Act.
Use the record to make the next decision
- Confirm the source item was eligible at the relevant time.
- Find the run, last confirmed stage and outbound attempts.
- Check the destination for the intended effect using the agreed reference or result ID.
- Classify the result as confirmed, rejected, accepted but pending, or still unknown.
- Record the resolution and owner of any remaining work.
Before relying on the record design, check whether a safe input rejection and a simulated lost response lead an operator to the right item and next action without exposing credentials or a full personal-data payload.

![How to choose the right idempotency key for an event: Keep the same idempotency key for retries of one action; use a new key for a new action.; Write: one [effect] for each [identity], e.g. one standard invoice per approved order.; Check stability, uniqueness, scope and lifetime; include tenant ID for tenant-scoped IDs. Choosing an idempotency key for an event](/covers/choosing-an-idempotency-key-for-an-event-640.webp?v=39704930)
