
Workflow Design
Part of Workflow rate limits and throughput
Planning queue capacity for a sudden event spike
Estimate backlog growth and drain time during an event spike, then check retention, worker capacity and deadlines.
Plan queue capacity for a spike by estimating arrivals, sustainable completions and how long the resulting backlog can wait. Storage alone is not useful capacity: an event may expire or become stale before anyone acts on it.
Calculate growth and drain
Choose an interval short enough to show the spike. For each interval, add arrivals to the existing backlog, subtract completed destination actions and floor the result at zero. A worker receiving a message is not a completed action. Include retries and slower record types when estimating sustainable completion capacity.
Suppose a hypothetical campaign produces 1,800 events evenly over five minutes, starting with an empty queue. That is 360 arrivals a minute. At 120 completed actions a minute, backlog grows by 240 a minute and reaches 1,200 after the spike.
If arrivals then fall to 30 a minute while capacity stays at 120, the backlog drains at 90 a minute. Clearing it takes about 13.3 minutes after the spike. The calculation assumes those rates remain steady and no other work competes for capacity; it is not a performance promise.
Check queue and worker limits
Review the selected service's retention, delivery semantics, in-flight limits and age metrics. Amazon SQS has finite, configurable retention and provides an approximate metric for the age of the oldest unprocessed message. Visible backlog alone omits work already held by consumers.
An SQS visibility timeout hides a received message while it is being processed. If the message is not deleted before the timeout expires, it can become visible again.
Set or extend the timeout to suit processing time. A short timeout can lead to repeat processing, while an unnecessarily long one can delay another attempt after a worker fails. Visibility timeout does not guarantee that a message is delivered only once.
The producer also needs a boundary: in Google Cloud, Pub/Sub publish requests can accumulate in client memory when publishing outruns transmission. Publisher flow control can limit outstanding requests by message count or bytes. If arrivals stay above completions, flow control alone cannot prevent backlog growth.
Key Queue Metrics for Spike Planning
- Message Retention (AWS SQS)
- Configurable up to 14 days
- Oldest Unprocessed Message Age
- Approximate metric available via CloudWatch
- Publisher Flow Control (Google Cloud Pub/Sub)
- Limits outstanding requests by message count or bytes
Protect the deadline
Set a maximum acceptable delay for each event class and compare it with predicted drain time and observed queue age. Route urgent work separately where routine backlog would block it. If applying a stale event late could overwrite newer state, hold it for a defined decision.
For a known launch or import, run a capacity exercise with representative event sizes and downstream API limits. Track arrivals, confirmed completions, throttling, visible and in-flight work, oldest age and exception volume.
During a real spike, these measures help determine whether to slow producers, pause lower-priority work or add consumers. More consumers help only while the destination can accept their combined output.


