
Observability
Workflow monitoring and observability
Monitor confirmed workflow outcomes, spot missing or overdue work and use execution signals to investigate the cause.
Monitor whether eligible business work reaches its intended destination, what remains unresolved and whether it finishes by the required time. Run metrics can locate trouble, but a completed run does not by itself confirm the business task was completed.
A paid order should produce a fulfilment request. The order system determines whether the order qualifies; the fulfilment system determines whether the request exists and what state it reached. Monitoring must account for an eligible order even if its trigger never starts a run.
Define the finish line
For each workflow, name the source record that creates an obligation, the destination state that fulfils it and the deadline. If the destination accepts work for later processing, track acceptance and final completion separately.
Give each eligible item one current outcome status: awaiting processing, in progress, confirmed complete, deliberately rejected under an agreed rule, or unresolved. Use outcome unknown when a write may have succeeded but its response was lost. Track overdue as a separate flag so an item can remain in its correct status while passing its deadline. Count ineligible items separately.
Question / Evidence to collect
- Did eligible work arise?
- Source identity and eligibility decision
- Did the workflow act?
- Run identity, stage reached and attempts
- Did the destination accept the request?
- Response or destination request ID
- Was the intended result achieved?
- Destination record or final state
- What still needs action?
- Open item, owner and age
Judge reliability by business behaviour
A workflow can be available and still be unreliable if it does not deliver the result people expect. OpenTelemetry illustrates this with a shopping-cart action that fails to add the selected item: service health alone does not establish user-facing reliability.
A service-level indicator (SLI) measures behaviour from the user’s perspective. A service-level objective (SLO) communicates reliability by connecting one or more indicators to business value.
Monitor the path, including its gaps
Compare eligible source items with confirmed destination outcomes for the same cohort. Show pending, overdue and unresolved items alongside those counts. Compare identities as well as totals: equal counts can conceal both a missing result and an unintended duplicate.
Use execution signals to investigate the gap: trigger activity, failed actions, retries and waiting work. Label timing precisely. Time to request acceptance differs from time to the completed destination action.
Platform views have limits. Azure Logic Apps documentation distinguishes trigger attempts from workflow instances: each time a trigger successfully fires, it creates an individual workflow instance.
AWS documentation says Step Functions CloudWatch metrics are delivered on a best-effort basis; completeness and timeliness are not guaranteed. These are product-specific limits; check the platform in use before treating its run chart as a business ledger.
Establish a baseline before interpreting change
AWS recommends measuring performance at different times and under different load conditions, then retaining historical monitoring data to compare current behaviour with normal patterns and anomalies.
AWS suggests monitoring execution starts and timeouts as baseline metrics for Step Functions, with activity starts and timeouts also relevant when Activities are used. Interpret them alongside the workflow’s eligible items and confirmed destination states.
Before monitoring, agree on the goal, resources, monitoring frequency and tools, as well as who performs the monitoring and who is notified when something goes wrong.
Make an affected item explainable
Metrics show a pattern; an operator needs the item behind it. Keep diagnostic context sufficient to connect an affected item with its workflow execution and destination evidence, where supported. OpenTelemetry describes metrics, logs and traces as different signals. A trace alone does not confirm an external business result.
Use signals to investigate unfamiliar behaviour
Observability supports questions about a system’s behaviour without requiring prior knowledge of its internal workings. OpenTelemetry describes it as a way to troubleshoot unfamiliar problems and ask why something is happening; that depends on applications emitting enough information to investigate without adding instrumentation during an incident.
A trace follows a request as it moves through distributed components, and its spans describe parts of that path. Logs are timestamped messages, but are not necessarily tied to a particular request or transaction. Together, these signals can help locate where behaviour diverged, while the destination’s business record remains the evidence that the intended outcome occurred.
Review outcome and execution signals
Review execution alerts against eligible source records and destination evidence: an alert is a prompt to investigate, not proof that a business outcome was missed.
Unresolved outcomes need an owner and business context; separate error-handling processes determine retry and alert routing.
Before relying on monitoring, trace representative eligible, ineligible, skipped-trigger, handled-failure and lost-response cases through the source, workflow and destination. Revisit the definitions when an event, destination operation or deadline changes.
In this guide
- Tracking successful outcomes rather than execution countsDefine eligible work, confirm destination results and calculate outcome and on-time rates without counting retries as extra successes.
- Recording enough context to diagnose a failed workflowCapture the IDs, attempted action, known outcome and safe error context needed to investigate a failed workflow.
- Building a workflow health dashboardLay out outcomes, backlog and execution signals with clear definitions, data freshness and routes to affected records.


