GKE architecture is often reduced to one choice: Autopilot or Standard. That is important and incomplete. A production cluster also depends on regional topology, node or compute-class strategy, networking, workload identity, upgrade policy, storage, scheduling, service exposure, observability, and the failure behavior of workloads. For Professional Cloud Architect, the strongest answer follows those dependencies rather than choosing a default and stopping.
Google recommends Autopilot for the most streamlined experience, with Google managing nodes and many infrastructure settings. Standard remains appropriate when teams need specific node-level control, and current GKE can also run Autopilot-mode workloads in Standard clusters. The architecture should therefore state which responsibilities the organization actually needs to own.
A strong baseline starts with Kubernetes fundamentals and then adds Google Cloud-specific operating choices. The cluster is not the application. It is a control plane and scheduling environment whose design should make application deployment, isolation, scaling, recovery, and upgrades predictable.
Choose the mode from required control
Autopilot removes much node administration and applies production-oriented defaults. Use Standard when a workload has a concrete need for node-pool management, host-level configuration, unsupported workload behavior, or another requirement that cannot be met within the managed model.
Write that requirement in the architecture decision record. ‘We may need more flexibility later’ is weak justification for permanently accepting more operational surface. A future workload can be evaluated when it exists.
Mode choice should also consider organizational ownership. Autopilot works well when application teams want Kubernetes APIs without owning node infrastructure. Standard may fit a platform team that deliberately provides specialized node pools, accelerators, kernel settings, or legacy compatibility as a shared internal platform.
Cluster location is an availability decision
Regional clusters replicate the control plane across multiple zones and can run workloads across zones, reducing dependence on one zone. Zonal clusters can be appropriate for lower-cost or nonproduction scenarios but create different failure exposure. Workload replicas and Pod topology still need to be designed; a regional control plane does not automatically spread every application safely.
Define what should survive a zone failure and test whether replica counts, topology constraints, persistent storage, and dependent services actually support that objective.
Regional design should include workload disruption budgets and replica placement. A multi-zone cluster cannot preserve service if all replicas of an application are scheduled into one zone or if the application has only one replica. Control-plane resilience and workload resilience are separate design tasks.
Topology policies should also consider planned maintenance and voluntary disruption, not only zone failure. PodDisruptionBudgets and spread constraints can preserve availability during upgrades only when replica counts and capacity are sufficient. A policy that forbids disruption without spare capacity can block maintenance instead of improving reliability.
Node and compute strategy affects scheduling behavior
Standard clusters can use multiple node pools with different machine types, accelerators, operating systems, labels, taints, and upgrade strategies. Autopilot provisions resources from workload specifications and current ComputeClass capabilities. In both cases, resource requests influence scheduling and cost.
Over-requesting CPU or memory wastes capacity; under-requesting creates contention and unstable scheduling. Use observed workload behavior and limits rather than copying generic resource settings between services.
Compute strategy should account for accelerators and specialized workloads. GPUs, local SSD, confidential or specialized machine types, and topology-sensitive applications can drive node-pool decisions in Standard. In Autopilot, verify current workload support and ComputeClass options before assuming the same control exists.
Node-pool sprawl in Standard is another risk. Creating a new pool for every workload characteristic can fragment capacity and complicate upgrades. Consolidate where requirements are compatible and introduce specialized pools only when hardware, security, isolation, or scheduling needs justify them.
Workload identity should replace node-wide credential assumptions
Pods need Google Cloud permissions for APIs and data services. Workload Identity Federation for GKE lets workloads use IAM identities without distributing long-lived service account keys and without treating the node’s identity as the security boundary for every Pod.
Map Kubernetes service accounts to the permissions each workload needs. The service-account model discussed in Google Cloud IAM applies here: authentication should identify the workload that acted, and compromise of one Pod should not automatically expose every permission held by the node or cluster.
Identity mapping should avoid broad Kubernetes service accounts shared across namespaces or applications. The cloud service account becomes the permission boundary for Google APIs, so one mapping reused by many workloads can turn a single Pod compromise into cross-application access.
Service-account permissions should also be reviewed when namespaces or workloads move between teams. A Kubernetes object can be redeployed quickly, but the Google Cloud IAM mapping may preserve access intended for the previous owner unless identity lifecycle is part of the migration checklist.
Network architecture extends beyond Services and Ingress
Pod and Service IP planning, VPC-native networking, subnets and secondary ranges, load balancing, Gateway or Ingress, egress, firewall policy, DNS, and network policy all affect application behavior. Address exhaustion or overlapping ranges can block growth just as effectively as a broken deployment.
Service-to-service communication may also introduce a service mesh or other policy layer. The internal discussion of Kubernetes service-mesh connectivity is useful when the application needs traffic identity, policy, or telemetry beyond basic Kubernetes Services.
IP planning should include Pod and Service ranges, future clusters, Shared VPC allocation, and hybrid connectivity. GKE address demand can be much larger than the number of nodes suggests, which is why subnet design needs platform-specific forecasts rather than generic VM counts.
Network design should also account for east-west service discovery and north-south ingress separately. A cluster may have healthy internal Services while external Gateway or load-balancer configuration is broken. Runbooks should identify which layer owns the failed path.
Security needs cluster and workload layers
Control-plane access, IAM, Kubernetes RBAC, workload identity, admission controls, Pod security settings, image provenance, secrets, network policy, and runtime monitoring protect different attack paths. One hardened layer does not compensate for broad privileges in another.
The principles in Kubernetes cluster security are valuable because Kubernetes security is about reducing the number of ways a compromised workload can become a compromised cluster or cloud identity. Default-deny thinking and least privilege should be applied to both Kubernetes and Google Cloud permissions.
Security posture should be validated after enabling third-party agents, privileged workloads, or host-level access. Exceptions to hardened defaults can be necessary, but every exception should identify the workload that needs it and the new attack path the team must monitor.
Supply-chain controls belong beside runtime controls. Image provenance, vulnerability scanning, signed artifacts, and trusted registries reduce the chance that a secure cluster runs an untrusted workload. Cluster hardening cannot compensate for deploying a malicious or unreviewed image.
Upgrades are part of the application lifecycle
GKE manages control-plane and node software differently depending on mode and configuration, but application teams still need to test workloads against Kubernetes version changes and deprecated APIs. Maintenance windows and release channels help schedule platform change; they do not guarantee application compatibility.
Keep a representative staging cluster or validation path, track deprecated APIs, and avoid blocking upgrades indefinitely with one unmaintained workload. Platform currency is a security and reliability requirement.
Upgrade strategy should consider both GKE release channels and application dependencies. A rapid channel can deliver features and fixes sooner but requires faster application validation. A more conservative channel reduces change frequency but does not remove the need to track API deprecations and supported versions.
Version-skew and add-on compatibility deserve explicit tests. Ingress controllers, CSI drivers, service meshes, policy engines, and observability agents can lag behind Kubernetes changes. The application may be compatible while a platform extension becomes the upgrade blocker.
Storage and state can constrain topology
Persistent workloads may depend on disk type, access mode, zone, backup, and recovery behavior. A StatefulSet can recreate Pods, but that does not make the application’s data globally available or automatically recoverable.
Document the stateful dependency separately from the Kubernetes object. Ask how data survives node loss, zone loss, cluster recreation, and operator error. Often the safest answer is a managed database or storage service rather than keeping critical state inside the cluster.
Persistent storage should be tested during rescheduling and zone failure. A Pod coming back is not enough; the volume must attach or the managed data service must remain reachable. Recovery exercises should include the stateful layer so Kubernetes health does not become a false indicator of application recovery.
A GKE design should be operable during failure
Review the architecture by removing one assumption at a time: a zone is unavailable, a node pool cannot scale, a Pod image is bad, DNS is slow, an identity mapping breaks, or an upgrade causes a compatibility issue. The owners and evidence for each failure should be known.
The Professional Cloud Architect certification is not testing whether an architect knows every GKE feature. It rewards the ability to choose boundaries and managed responsibility deliberately. Autopilot versus Standard is one choice inside that larger operating system.
Operational evidence should include cluster events, workload logs, control-plane or API signals, node and Pod resource behavior, and cloud-service dependencies. The fastest GKE incidents are resolved by teams that can distinguish scheduling failure, network failure, identity failure, and application failure quickly.
SLOs should distinguish platform health from application health. The Kubernetes API can be available while a business service is down because dependencies, configuration, or data are unhealthy. Monitor user-facing service objectives in addition to cluster-level metrics.
A platform review should also include ownership of cluster-wide extensions. Ingress controllers, policy engines, service-mesh components, and observability agents can affect every namespace; they deserve versioning and change control comparable to other shared infrastructure.