Observability
Part of Workflow rate limits and throughput
Measuring the delay between trigger and completed action
Measure workflow delay from source event to confirmed destination action, including queue and retry time.
Measure workflow delay from the source event to the confirmed destination action. A successful run or API acknowledgement can be an earlier milestone if work continues afterwards. Define the finish line before calculating delay.
Record milestones for each event
Join records across systems with a stable event ID or correlation key. Capture the source event time, workflow receipt, queue entry, processing attempts, outbound requests and confirmed destination completion.
Store timestamps with a consistent time basis and time zone information. Check clock synchronisation where systems record their own times.
Keep provider request and message IDs alongside the business event ID, but do not treat them as the same identifier.
For a completed event, total delay is destination completion time minus source event time. Break it into trigger delivery, queue wait, processing, retry wait and destination confirmation where the timestamps allow it. Avoid adding overlapping intervals twice.
If the destination only acknowledges receipt and processes work asynchronously, measure the later completion point. If that point cannot be observed, call the metric time to acceptance.
Interpret platform timestamps carefully
Amazon SQS can return ApproximateFirstReceiveTimestamp, which records when a message was first received from the queue. It can help estimate the wait before first receipt, but it does not cover later re-queues or destination work. SQS also reports ApproximateAgeOfOldestMessage, an approximate queue-level signal rather than an event's end-to-end delay.
OpenTelemetry's messaging conventions distinguish client operation duration from message processing duration. Propagated message context can correlate producer and consumer traces. Those conventions are still marked as development, and a trace needs the application's actual completion point to establish the business outcome.
Show slow and unfinished work
Report a distribution of completed-event delays, including a median and a high percentile chosen for the business deadline. Alongside it, show the share completed within the deadline and how many events remain pending beyond it or failed. Otherwise the completed-delay chart can appear to improve when the slowest events never finish.
Split results by workflow, destination, event class and time window. Compare queue and outbound waits to identify the slow stage. For example, an order update may reach the trigger promptly, wait in a queue during a promotion and then wait through API throttling.
Before relying on the measure, trace a safe sample event from source to destination and compare recorded milestones with the observed destination state. Check how failed and retried events appear, especially after changing the trigger or destination integration.



