Amazon AWS SAA-C03: Aurora Global Database

Amazon Aurora Global Database extends one Aurora database across multiple AWS Regions by using a primary cluster for writes and read-only secondary clusters for low-latency reads and disaster recovery. Replication uses Aurora’s storage layer rather than ordinary database-engine replication, which helps keep cross-Region replication latency low. The important operational point is that the architecture is still single-primary: changing the writer Region is a controlled promotion event.

Within AWS Architecture and Operations, Aurora Global Database is a business-continuity design. The service provides the cross-Region database mechanism, but the application still needs connection strategy, failover runbooks, DNS or endpoint behavior, lag monitoring, and recovery testing.

The existing Aurora architecture under real load article provides the local-cluster performance context.

Switchovers are for planned events and target zero data loss

Aurora calls the planned operation a switchover. AWS previously called this a managed planned failover. During a switchover, all clusters and dependent services are expected to be healthy, and Aurora waits for the selected secondary to synchronize with the primary before changing roles.

Because the target is synchronized first, the switchover RPO is zero. The old primary becomes read-only, the chosen secondary promotes a writer, and the global topology remains intact.

This makes switchovers appropriate for maintenance, regional rotation exercises, follow-the-sun write placement, or failback after a disaster-recovery event.

Failover is for an unplanned primary-Region outage

A managed failover is intended for a true primary-Region outage or service-level failure. Aurora promotes a chosen secondary without waiting for the original primary to synchronize because the original region may be unreachable.

That means the RPO can be non-zero. Any transactions that had not replicated to the selected secondary before the outage can be lost. The database remains transactionally consistent, but “consistent” does not mean “contains every acknowledged write from the failed Region.”

Applications with strict recovery requirements should therefore monitor replication lag and define how much data loss is acceptable before an incident occurs.

Managed failover preserves topology better than manual promotion

Aurora supports managed and manual failover approaches. Managed failover is the recommended disaster-recovery method when the global database configuration supports it. When the old primary recovers, Aurora can reintroduce it as a secondary and preserve the global cluster topology.

Manual failover remains relevant when managed failover cannot be used, such as incompatible engine versions. Manual procedures require more operational care and should be documented and rehearsed before they are needed.

The recovery plan should not begin with engineers reading documentation during a regional outage.

Engine-version compatibility affects managed cross-Region operations

Switchovers and managed failovers depend on compatibility between the primary and secondary engine versions. AWS documents which Aurora MySQL and Aurora PostgreSQL versions support cross-Region promotion with matching or compatible patch levels.

That means patch management is part of disaster recovery. A secondary that lags behind in engine version can weaken the failover plan even though replication is otherwise healthy.

Upgrade standards should preserve a failover-compatible topology or document the temporary period when managed promotion is unavailable.

Global writer endpoints reduce application reconfiguration

Aurora Global Database provides a global writer endpoint that tracks the writer in the current primary cluster. Applications that use the global endpoint can avoid hard-coding one regional cluster writer endpoint and can reconnect after the primary Region changes.

This still requires connection management. Existing database connections do not teleport to the new Region. Applications need retry and reconnect logic that can tolerate the short unavailability during switchover or failover.

Connection pools should be tested so stale connections do not keep pointing at the former primary after the topology changes.

Write forwarding enables secondary-region applications without multi-primary semantics

Aurora Global Database can support write forwarding on supported engine versions. An application connected to a secondary cluster can issue supported write statements that Aurora forwards to the primary writer.

Read consistency can be configured so the secondary waits for different degrees of replication visibility. Aurora PostgreSQL exposes consistency modes such as SESSION, EVENTUAL, GLOBAL, and OFF; Aurora MySQL exposes its own corresponding session-level controls and constraints.

Write forwarding reduces application routing complexity for some global designs, but writes still execute in the primary Region. Cross-Region network latency and primary connection capacity therefore remain part of the write path.

Read scaling should not hide stale-data expectations

Secondary Regions are attractive because they bring reads closer to global users. Those reads are fed by asynchronous cross-Region replication, so they can lag the primary.

Applications should decide which data can tolerate eventual consistency and which operations require stronger guarantees. A global catalog page may tolerate slight lag, while a just-completed financial transaction may not.

If the user experience mixes local reads with forwarded writes, consistency settings should be tested explicitly so the application does not appear to “lose” a write that is merely waiting to replicate.

CloudWatch lag metrics should be part of the failover runbook

AWS recommends monitoring secondary replication lag and using that information when selecting a recovery target. The runbook should identify which metrics and thresholds operators will use during a real outage.

The broader multi-Region disaster recovery article is relevant because a database failover is only one component of application recovery. Network routes, secrets, application capacity, queues, and caches must also work in the recovery Region.

Failover tests should include the application path, not only the database control plane.

Global database is useful when the recovery objective justifies the complexity

Aurora Global Database introduces cross-Region cost, replication, version-management, connection, and testing responsibilities. It should be used because the application needs global read locality, regional continuity, or both—not because multi-Region architecture looks more resilient on a diagram.

The design is mature when the organization knows its RPO and RTO, has a tested promotion mechanism, understands the difference between planned switchover and unplanned failover, and can explain exactly how the application reconnects to the new primary.

Secondary-region capacity should be production-ready before it is considered a failover target. A headless secondary without a DB instance cannot be promoted until an instance is added, and an undersized secondary can meet replication requirements while still failing the application’s post-promotion load. Recovery planning should include writer instance class, reader capacity, connection limits, parameter groups, and the surrounding application tier.

RPO and RTO should be measured in exercises rather than copied from service descriptions. Replication lag during normal traffic, promotion duration, application reconnect time, DNS or endpoint propagation, cache warm-up, and downstream dependency recovery all contribute to the observed outcome. The database control plane can complete quickly while the application remains unavailable because clients are still using stale connections.

Write-forwarding designs need capacity protection on the primary writer. Forwarded sessions consume connection resources on the primary, and AWS exposes engine-specific controls for how much of the writer’s connection capacity can be used for forwarded activity. A secondary-region application can therefore increase load on the primary even when its local reads remain regional.

Consistency modes should be chosen from user expectations. EVENTUAL consistency minimizes waiting but can expose replication lag after a write. SESSION consistency can ensure a session sees its own forwarded changes, while stronger global consistency waits for more complete visibility and can increase read latency. The product should define which interactions require read-after-write behavior rather than applying the strongest setting everywhere.

Failback deserves its own rehearsal. After an unplanned failover, the new primary may run in the secondary Region for hours or days while the original Region is repaired. Returning writes to the original Region should use a planned switchover after the topology is healthy, not a rushed reverse failover during an ongoing incident.

Backups and snapshots remain necessary even with a global database. Cross-Region replication protects availability and reduces recovery time, but logical corruption or an erroneous application write can replicate to every Region. Recovery architecture should therefore combine global replication with backup retention, point-in-time recovery, and tested restore procedures.

Operational exercises should include application write quiescing for planned switchovers. Even though Aurora synchronizes the target secondary, upstream applications, background jobs, and integration workers should be coordinated so they do not create avoidable connection errors during the brief role transition. A controlled maintenance window can validate how quickly the full application stack follows the new writer.

Read endpoints should also be reviewed after promotion. Applications that intentionally used a local secondary reader for latency may need different routing once that Region becomes primary. Monitoring and configuration should distinguish “nearest reader” logic from “writer” logic so a topology change does not accidentally send read-heavy traffic to an unexpected Region.

Cost planning should include replicated storage and secondary compute. A global database kept only for disaster recovery can still incur continuous cross-Region and instance cost, especially when secondaries are sized to take production traffic immediately. That cost is part of the resilience decision and should be allocated to the service that benefits from the lower RTO.

The database team should also document unsupported or restricted features for the selected engine/version combination. Global write forwarding, switchover behavior, patch compatibility, and local write forwarding can differ between Aurora MySQL and Aurora PostgreSQL, so architecture standards should be engine-specific rather than one generic “Aurora Global” checklist.

Leave a Reply

How It Works

img
Step 1. Choose Exam
on ExamLabs
Download IT Exams Questions & Answers
img
Step 2. Open Exam with
Avanset Exam Simulator
Press here to download VCE Exam Simulator that simulates real exam environment
img
Step 3. Study
& Pass
IT Exams Anywhere, Anytime!