An AWS Lambda function consuming Amazon SQS messages often receives several records in one invocation. One invalid message can cause the whole batch to be retried when the function throws, even if earlier records were processed successfully. That behavior creates duplicate work, amplifies downstream traffic, and can hide a poison message within a busy queue. Partial batch responses let a consumer identify individual records that failed, but they only work when the function and event source mapping use the contract correctly.
This design question is more than an optimization. It determines whether a consumer preserves message order, whether business operations tolerate retries, and when an unprocessable record reaches a dead-letter queue. The correct goal is not to make every batch appear successful. It is to preserve a precise account of which records committed their side effects and which ones remain eligible for another attempt.
Understand the default SQS retry boundary
Lambda polls the queue and makes records available to the function according to batch size, batching window, and available concurrency. An SQS message is not deleted merely because it was delivered. Successful processing leads the Lambda integration to remove the message from the queue; an unsuccessful batch can become visible again when the visibility timeout expires.
If the handler raises an exception after processing some records, the integration generally treats the invocation as failed. Completed records may return with the failed record because the batch was not acknowledged. This is not evidence that SQS lost its place; it reflects a batch-level acknowledgment boundary. Downstream handlers must tolerate repeated deliveries whenever a retry or delivery race occurs.
A queue’s visibility timeout should account for function execution, retry behavior, and batching. Setting it too short can expose messages before processing finishes, creating concurrent handling of the same work. Setting it excessively long delays recovery from failed consumers. Test observed runtimes rather than deriving the timeout from the nominal handler duration alone.
Enable the partial response contract explicitly
For an SQS event source mapping, configure ReportBatchItemFailures in FunctionResponseTypes. Without that setting, a handler that returns a JSON object listing failed message IDs may not change how the integration handles the batch. The mapping configuration is therefore part of the application contract and should be versioned with deployment infrastructure.
The handler returns a batchItemFailures collection with itemIdentifier values equal to the failed messages’ identifiers. Successfully processed messages are omitted. Use message IDs from the input event rather than application identifiers or receipt handles. A malformed response or an uncaught function exception can defeat the intended record-level acknowledgment behavior.
A batch response should only mark a record successful after its required side effect has committed. If a handler initiates an asynchronous write and returns before checking the result, the message can be removed while the write later fails. Conversely, reporting a committed record as failed will replay it, so idempotency remains necessary even after partial responses are implemented.
Handle standard and FIFO queues differently
For a standard queue, a handler can process records independently when the business operation has no cross-record ordering requirement. Collect failures and return their message IDs. However, individual failures should not silently stop checks of other independent records; whether to continue should follow the application’s throughput and downstream load policy.
FIFO queues require stricter treatment of message order. AWS recommends stopping processing after the first failure in the relevant ordering sequence and returning failed and unprocessed records in the response. If later messages from the same message group are acknowledged while an earlier one failed, the consumer can violate the business ordering that FIFO was intended to protect.
Message-group concurrency also matters. A long-running or repeatedly failing record can block progress in its group while other groups continue. Record the group identifier and sequence context in diagnostics so an apparently healthy overall queue does not conceal a particular customer or entity whose ordered updates have stopped.
For an order-confirmation consumer, store a durable operation key with the committed order transition. Suppose a record triggers payment confirmation and the response times out after the payment gateway commits. The handler should query the gateway by that key before making another payment request. If the gateway confirms success, the SQS record can complete without repeating the charge; if status is uncertain, keep the record retriable and raise an actionable alert. This protects against the common failure mode where the Lambda runtime sees an exception but the customer’s account has already been affected.
Build idempotency around business outcomes
At-least-once message delivery means a handler must recognize already-applied work. An order event can carry an operation ID that is stored with the completed order transition; a payment action might require a durable idempotency key provided to the payment gateway. Checking merely whether a Lambda invocation has seen the SQS message ID is insufficient if that ID is regenerated when the business request is retried upstream.
Place the deduplication check near the committed effect. A database write and an idempotency record in separate nontransactional systems can fail between operations, leaving an ambiguous result. Use transactions where possible or design recovery queries that can determine the actual state after a timeout. A retry should converge toward the intended outcome rather than add another identical side effect.
The Lambda event-driven pattern includes event source behavior, consumer state, and downstream failure handling. Partial batch acknowledgments reduce redundant work; they cannot replace safe operation identifiers, business constraints, or monitoring of failed external API calls.
Distinguish transient failures from poison records
A temporary database timeout may succeed on a later delivery, while a schema violation or impossible business transition can fail on every attempt. Treating both as unlimited retries wastes capacity and can prevent useful work from progressing. Classify errors, record machine-readable reasons, and decide which conditions should be retried automatically.
Configure source queue redrive policy and maxReceiveCount with the actual retry and visibility behavior in mind. A dead-letter queue preserves failed messages for triage, but it does not repair the originating data or guarantee that a human can safely replay it. Preserve enough context to locate the producing system and the intended business operation without leaking sensitive payloads into broad-access logs.
When a poison record enters a DLQ, create an explicit resolution path: correct the producer, migrate an outdated payload, apply a justified business exception, or mark it permanently rejected. Simply moving the same invalid record back to the source queue can recreate the incident. Link DLQ redrive to a versioned investigation outcome and replay safeguards.
Control concurrency and downstream pressure
Enabling partial responses can change the poller’s response to errors; with ReportBatchItemFailures, Lambda does not scale down message polling merely because individual records fail. This helps throughput when failures are isolated, but it can intensify a downstream outage if many records keep retrying. Concurrency and retry policies need to match the capacity of dependent services.
Set reserved concurrency, event-source maximum concurrency where supported, and downstream request limits deliberately. A queue backlog can grow during a dependency outage without requiring the consumer to hammer the unavailable system. Use backoff in client calls and allow messages to return to the queue instead of holding executions in long in-function sleep loops.
A Lambda SQS consumer needs its failure-response format, visibility timeout, concurrency, and retry strategy to agree; DVA-C02 developer decisions must prevent one bad record from reprocessing successful work. A correct design balances throughput, retry behavior, visibility timeout, idempotency, and batch response format. Tuning one parameter in isolation can transform a small number of bad records into a wider operational failure.
Instrument per-record outcomes and reconciliation
Capture input message ID, application operation ID, message-group information where relevant, attempt context, and outcome category. Logs should distinguish committed success, retryable failure, permanent rejection, and ambiguous external effect. Counting successful Lambda invocations alone can conceal large numbers of records returned for retry within otherwise successful batches.
Monitor queue age, approximate visible and not-visible message counts, DLQ growth, handler error rates, and per-record failure classifications. Alert when the oldest message age grows rather than only when Lambda errors spike. A handler that routinely returns a partial failure response may produce few invocation-level errors while one class of messages remains undelivered.
Reconcile consumer output with source-system expectations. For a financial adjustment queue, compare accepted adjustment IDs against committed ledger changes. For event enrichment, compare expected keys with populated derived records. This catches silent loss caused by incorrect success reporting and distinguishes genuine backlog from duplicates already processed.
Test a batch in which one poison record appears ahead of nine valid records, then repeat the scenario with the poison record at the end. In a standard queue, the desired failure identifiers may be identical in kind but differ in count depending on independent successes. In a FIFO group, the early failure should also return subsequent unprocessed identifiers to preserve order. Compare actual deletion and retry behavior after visibility timeout, not just a mocked handler response. Confirm the DLQ receives only records that exhausted their permitted delivery attempts and retains enough investigation metadata for a safe manual decision.
Test failure sequences before production
Create batches that include success followed by failure, failure followed by success, multiple independent failures, a dependency outage, and a failure after a downstream commit. Confirm the handler returns exactly the identifiers that should retry. For FIFO queues, confirm later records in the affected sequence are not prematurely acknowledged.
Include deployment regression tests for mapping configuration. A code-only unit test cannot detect an event source mapping that lost ReportBatchItemFailures during infrastructure replacement. Similarly, test batch size and response shape after runtime upgrades or framework middleware changes; a utility wrapper may handle responses differently from the raw handler.
Partial batch response is effective when acknowledgment represents durable business completion. The technical setting is simple, but its safety depends on queue semantics, failure classification, idempotency, capacity, and observable reconciliation. Those factors determine whether a failed record stays isolated or repeatedly disrupts an entire event-driven workload.