Batching records without losing failures: Use stable identifiers to track each record in a batch.; AWS SQS allows up to 10 entries per batch with separate success/failure results.; Reconcile outcomes by ID, not position, and retry only failed entries.
Image: Workflow Automation Guide

Data Mapping

Part of Workflow rate limits and throughput

Batching records without losing individual failures

Use batches while tracking each record’s success, failure or unknown outcome and retrying only appropriate entries.

Batching saves requests only if the workflow can still account for every record. Give each input a stable identifier, map it to the submitted batch entry and reconcile the result before closing that stage of work. An HTTP success for a batch does not necessarily mean that every entry succeeded—or that a later destination action is complete.

Check the operation’s contract

Find the entry-count and payload limits, whether the batch can partly succeed, and how its response identifies individual results. Check whether order matters and how long a record waits for the batch to fill. Larger batches can improve call efficiency while adding wait time.

Amazon SQS SendMessageBatch accepts up to ten entries and returns separate Successful and Failed results. AWS warns that entry failures can appear even when the request returns HTTP 200.

A successful entry means its message was accepted into SQS; it does not prove that a consumer completed the later business action. Other batch operations may have different semantics.

AWS SQS SendMessageBatch: Success vs Failure Handling

Maximum entries per batch
10
Response includes individual results
Yes (Successful and Failed lists)
HTTP 200 means all succeeded?
No – some may have failed internally
Partial success supported?
Yes

Reconcile each record

Keep the source event ID, batch entry ID, operation and attempt number. After a response, classify each entry as confirmed success for that operation, confirmed retryable failure, confirmed permanent failure or unknown outcome. Preserve the provider error code and request ID where available. Match results by the documented entry identifier, not by array position unless positional matching is guaranteed.

In a hypothetical eight-record batch, six entries succeed, one fails validation and one is throttled. Close the six for this operation, send the invalid record for correction and retry only the throttled record within the API's limits.

If the connection drops before the response arrives, all eight outcomes may be unknown. Check the destination or use a supported idempotency mechanism before resending writes that might have succeeded.

Key Metrics for Batching Record Workflows

Max batch size (SQS)
10 entries
Typical response time (batch)
Milliseconds to seconds
Idempotency required for retries?
Yes – to prevent duplicate processing

Account for consumer failures

A queue can also deliver a batch of messages to one worker. With an AWS Lambda SQS event source, a failed invocation normally makes the whole batch visible again, including messages the function processed successfully.

AWS supports partial batch responses through ReportBatchItemFailures, but the event source mapping must be configured and the function must return failed message IDs. An uncaught exception still fails the whole batch.

Acknowledge or delete source work according to the source system's contract and only after the downstream action required for that stage is confirmed. Reconcile submitted, successful, failed and unknown counts before closing each batch. Monitor ageing failed records separately so high batch throughput does not hide missing actions.

Partial Batch Failures in AWS Lambda SQS Event Sources

  • ProsAllows successful messages to be acknowledged even if others fail; reduces data loss risk when using `ReportBatchItemFailures`
  • ConsRequires explicit configuration and function code to return failed message IDs; uncaught exceptions still cause full batch rollback

More from Data Mapping

Observability

Measuring the delay between trigger and completed action

Measure workflow delay from source event to confirmed destination action, including queue and retry time.