Google Cloud Architect: Pub/Sub Delivery Semantics

Google Cloud Pub/Sub is often described as a decoupling service: publishers send messages to a topic, subscribers consume them independently, and the platform buffers traffic between the two. That description is correct but incomplete. Reliable designs depend on delivery semantics—what can be delivered more than once, what can arrive out of order, when an acknowledgment is final, and how the system behaves when a subscriber fails.

For Professional Cloud Architect scenarios, the key lesson is that messaging reliability is shared between the service and the application. Pub/Sub provides durable delivery mechanisms, retries, dead-letter handling, optional ordering, and exactly-once support for eligible subscriptions, but consumers still need correct acknowledgment and side-effect behavior.

By default, Pub/Sub provides at-least-once delivery and no ordering guarantee. That default should shape the consumer from the first line of code rather than being treated as an edge case discovered after duplicate business events appear.

At-least-once means duplicates are part of normal design

A subscriber can receive a message more than once if an acknowledgment does not complete before the deadline or the client negatively acknowledges the message. Network failures can also create ambiguity about whether processing succeeded.

The safest consumer assumes repetition is possible. Use idempotency keys, deduplication state, conditional writes, or business operations that are naturally safe to repeat. A payment capture, inventory decrement, or email send deserves different duplicate-handling logic than a metrics update.

The broader design pattern in event-driven Pub/Sub systems is useful: decoupling shifts some coordination from synchronous calls into message state. That independence is powerful only when the consumer can tolerate redelivery.

Acknowledgment is a reliability boundary

Acknowledging too early risks message loss from the application’s perspective because the service believes work is complete before the durable side effect has finished. Acknowledging too late can create unnecessary redelivery and duplicate processing.

Place the acknowledgment after the point at which the required work is durably complete. If processing takes longer than the acknowledgment deadline, use supported client-library lease extension behavior or redesign the work into smaller units.

The correct boundary is business-specific. If the consumer writes to a transactional store and then emits another event, decide which operation defines completion and how the second operation is recovered if it fails.

Retry policy should match the failure being retried

Pub/Sub can retry unacknowledged messages immediately or apply exponential backoff within supported bounds. Immediate retry is suitable when failures are brief and the dependency is likely to recover quickly. Backoff is safer when repeated attempts would overload an unhealthy downstream system.

Do not use message retry as a scheduling mechanism. A subscriber that intentionally fails work for hours to create a delay turns the error path into a control plane and makes operational signals misleading.

This mirrors the principle behind event-driven retry design: a retry should respond to transient failure, not compensate for missing workflow logic.

Dead-letter topics are an operational escape valve

Messages that repeatedly fail should not block healthy traffic forever. A dead-letter topic provides a place for undeliverable messages after a configured range of attempts, allowing the main subscription to continue making progress.

Treat dead-letter messages as unresolved business work, not garbage. Capture enough context to diagnose why they failed, alert on dead-letter volume, and define a replay or remediation process.

A growing dead-letter queue can indicate schema drift, dependency failure, invalid data, authorization changes, or a software regression. Operators should be able to distinguish those causes without manually opening random messages in production.

Ordering is scoped to ordering keys, not the entire topic

Pub/Sub message ordering can preserve order for messages with the same ordering key when the feature is enabled and the publishing conditions are met. It does not turn a distributed topic into one global serial queue.

Choose ordering keys around the entity that genuinely requires sequence: an account, device, order, or workflow instance. A single key for all events destroys parallelism. Too many arbitrary keys can complicate reasoning without providing useful guarantees.

When ordered delivery is enabled, an earlier unacknowledged message can hold back later messages for the same key. That means a poison message can create localized head-of-line blocking even while unrelated keys continue to progress.

Exactly-once reduces duplicate delivery but not application complexity to zero

Pub/Sub exactly-once delivery is available for supported pull subscriptions and gives subscribers stronger acknowledgment semantics. After successful acknowledgment, the service does not redeliver the message, and clients can determine whether acknowledgment succeeded.

That is useful, but “exactly once” should not be interpreted as a magical end-to-end transaction across every system the consumer touches. A database commit can still succeed before the process crashes, leaving application state that must be reconciled with message acknowledgment.

Use exactly-once where the stronger service guarantee reduces complexity or risk, but keep business operations idempotent where practical. The service controls message delivery; the application controls its side effects.

Push, pull, and export subscriptions create different failure surfaces

Pull subscribers control message retrieval and acknowledgment directly, which makes them a strong fit for workers that need custom concurrency, batching, or exactly-once delivery. Push subscriptions move delivery over HTTP and depend on the endpoint’s response behavior. Export subscriptions deliver into supported storage or analytics destinations with their own operational model.

Choose based on the consumer contract. An HTTP application that already scales reliably may prefer push. A high-throughput worker with custom backpressure logic may prefer pull. A data-retention pipeline may use an export subscription.

The comparison with other messaging patterns is helpful because “message service” is not one universal operating model. Delivery type changes how failures, scaling, and acknowledgments are handled.

Monitoring should focus on age and progress, not only message count

A subscription can contain many messages and still be healthy if consumers are keeping pace. Conversely, a small backlog can be serious if its oldest message is far beyond the service-level objective.

Track backlog size, oldest unacknowledged message age, acknowledgment latency, delivery attempts, dead-letter volume, and subscriber errors. For exactly-once subscriptions, monitor acknowledgment-related failures and expired acknowledgment deadlines.

Correlate Pub/Sub metrics with the consumer’s own latency and error data. A backlog spike caused by an intentional maintenance window is different from a backlog that grows because a code release broke deserialization.

Oldest-unacknowledged-message age is often more actionable than raw backlog size. A large subscription can be healthy if consumers continuously drain messages within the service objective, while a small backlog containing very old messages can indicate a poison message, ordered-key blockage, or repeatedly failing dependency. Alert on age and growth rate together so operators can distinguish normal bursts from stalled progress.

Monitor retry and dead-letter behavior as a separate reliability signal. A dead-letter topic that never receives messages may mean the system is healthy, but it can also mean forwarding permissions or dead-letter configuration are wrong. Conversely, a sudden rise in dead-letter volume shows that retries are exhausting without resolving the underlying failure. Track the original subscription, delivery attempts where available, and consumer error categories so the dead-letter stream becomes diagnosable rather than a second backlog.

Ordering changes the interpretation of those metrics. When ordered delivery is enabled, one slow or failing message can delay later messages with the same ordering key even while other keys continue progressing. Aggregate throughput can therefore look normal while a business entity is effectively stuck. Operational dashboards should expose lag by key or by a meaningful shard where feasible, and application logs should preserve ordering-key context so responders can identify whether the problem is global, partition-specific, or isolated to a single sequence.

Event consumers need explicit side-effect strategy

Many duplicate problems are really side-effect problems. The same message can be processed twice without harm if the destination write is conditional on a unique event ID. The same message can be catastrophic if it initiates an irreversible external action twice.

For high-risk actions, record processing state in a durable store before or alongside the side effect. For lower-risk analytics updates, a naturally idempotent upsert may be sufficient. The level of protection should reflect the business impact.

A real-time event pattern such as event notifications reinforces the point: reliable architecture depends less on hoping for perfect delivery and more on designing the consumer so imperfect conditions remain safe.

Delivery semantics should be written into the architecture contract

Document the expected subscription type, ordering behavior, retry policy, dead-letter handling, acknowledgment boundary, and duplicate strategy for every important event flow. Without that contract, teams make inconsistent assumptions and failures become difficult to reproduce.

Review the contract when traffic volume or business criticality changes. A low-volume notification flow may not need exactly-once delivery, while a financial or inventory workflow may justify stronger semantics and more operational controls.

Across Google Cloud, Pub/Sub is strongest when treated as durable infrastructure with explicit application responsibilities. The service can deliver reliably, but correctness emerges only when the subscriber is designed for redelivery, ordering constraints, partial failure, and the reality that distributed systems do not complete every step at the same instant.

Consumers should be tested for duplicates, redelivery, ordering assumptions, and retry side effects under realistic failure. Pub/Sub delivery semantics become an application property only when handlers are idempotent and state transitions remain correct after the same message arrives more than once.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!