Hub-and-spoke networking is attractive because the first diagram is clean: shared connectivity and controls sit in a hub, application networks live in spokes, and peering connects the pieces. Problems usually appear later, when the environment adds regions, subscriptions, overlapping ownership, private endpoints, inspection requirements, hybrid routes, and hundreds of spokes.
The current AZ-305 exam expects architects to reason about infrastructure design, not merely recognize topology shapes. The important question is how the topology behaves when part of the system is unhealthy or when scale changes assumptions that were invisible at the beginning.
A hub-and-spoke design is therefore best reviewed through failure and ownership: what is centralized, what depends on it, which routes are stateful, who can change them, and how workloads behave when the hub cannot provide a shared service.
The hub becomes a dependency faster than teams realize
Centralization creates efficiency by concentrating shared services such as hybrid connectivity, firewalls, DNS, routing, and management. It also creates shared failure domains. If every spoke requires a firewall in one hub region to reach critical services, that firewall path is part of every workload’s reliability model.
This does not mean centralization is wrong. It means the reliability target of the shared component must be at least as carefully designed as the workloads that depend on it. A resilient application cannot meet its objective if its only egress, name resolution, or on-premises route depends on a weaker platform path.
The internal explanation of hub-and-spoke topology is a useful starting point; architecture work begins when the team traces those dependencies under failure.
Route intent and route reality can diverge
User-defined routes, BGP propagation, virtual network peering, VPN or ExpressRoute gateways, Azure Firewall, and network virtual appliances can all influence the effective path. Teams often understand the intended route but not the full set of mechanisms that can change it.
This becomes dangerous during partial change. A new route table is associated with some subnets but not others. BGP introduces a more specific prefix. A forced-tunneling design changes default routes. A spoke is peered to the wrong hub. Traffic still moves, but not through the inspection or failover path the architecture assumes.
Troubleshooting should therefore use effective routes and observed flow, not only diagrams. The question is not “What should this packet do?” but “Which decision actually caused this packet to take this path?”
Peering scale exposes management weaknesses
Manual peering works when there are ten spokes. It becomes risky when there are hundreds. Every new virtual network can require correct peering, route associations, DNS settings, network-security configuration, and registration with central services.
At scale, the topology needs automation and policy. The platform team should know how spokes are discovered, how unauthorized peerings are prevented, how route requirements are enforced, and how address allocations are checked before deployment.
That is where the operational skills represented by AZ-104 intersect architecture. A design that depends on perfect manual configuration is not a scalable design even if each individual configuration is valid.
Address space is an architectural resource
Overlapping address space can block peering, complicate hybrid connectivity, and create painful migration work. Early environments often allocate large prefixes casually because address scarcity does not feel immediate. Multi-region growth, acquisitions, partner networks, and on-premises connectivity make those choices expensive later.
A durable plan considers the entire corporate address estate, expected regions, spoke sizes, expansion space, non-Azure networks, and the possibility that some workloads may later need connectivity that was not initially planned.
Address efficiency is not about squeezing every subnet. It is about avoiding fragmentation and overlap while preserving enough structure that operators can understand what a prefix represents.
DNS can fail even when the network path is open
Private endpoints, hybrid name resolution, custom DNS servers, Azure-provided DNS, and conditional forwarding create a control plane for reachability that is easy to overlook. A route can be correct while the application still fails because the name resolves to an unexpected public address, an unreachable private address, or nothing at all.
Centralized DNS architecture should therefore be reviewed like centralized routing. Which resolvers do spokes use? How are private DNS zones linked? How are on-premises domains forwarded? What happens if the central resolver is unavailable? Which team owns changes?
Network troubleshooting that stops after a successful TCP path can miss the real dependency. In cloud architectures, name resolution is often part of the application topology.
Inspection requirements can create asymmetric paths
Stateful firewalls expect return traffic to traverse a compatible path. Hub-and-spoke designs can accidentally create asymmetry when traffic enters through one path and returns through another, particularly with multiple hubs, gateways, appliances, or direct spoke connections.
The architecture should make the inspection boundary explicit. Which flows must traverse the firewall? Are spoke-to-spoke flows inspected? Is internet egress centralized? How do private endpoints behave? What happens during regional failover?
The answers affect both security and reliability. Forcing every flow through a single inspection point simplifies policy but increases dependency and latency. Allowing direct paths can improve performance but weakens central enforcement unless compensating controls exist.
Multi-region design changes the topology question
A single hub can be reasonable in one region. Multi-region workloads usually need a clearer model: regional hubs, global connectivity, shared or distributed security controls, and a plan for how traffic shifts when a region is impaired.
The correct design depends on latency, hybrid entry points, stateful inspection, data location, and recovery objectives. A regional outage should not require every surviving workload to hairpin through the failed region’s networking stack.
Architects should test the topology with a region-loss scenario before production. If the answer depends on undocumented route changes performed during the outage, the recovery design is incomplete.
Operational ownership is part of the network design
Central networking creates a platform product. Workload teams need a clear contract: how to request connectivity, what address ranges are allowed, which ports can be opened, how DNS is registered, what telemetry is available, how incidents are escalated, and which changes require central review.
Without that contract, the hub becomes an organizational bottleneck. Teams work around controls, add local appliances, or create unofficial peerings because the approved path is too slow. The technical architecture then diverges from the governance model.
The broader Microsoft platform offers managed options and automation tooling, but ownership still has to be designed by the organization.
Use a simple mental model: path, state, dependency, owner
When a hub-and-spoke environment misbehaves, start with four questions. What exact path should the flow take? Which components on that path maintain state? Which shared dependencies must be healthy? Who owns each component and can change it?
This model works for routine troubleshooting and architecture review. It exposes route assumptions, asymmetric inspection, DNS dependencies, central bottlenecks, and gaps in operational responsibility.
Hub-and-spoke networking scales well when centralization is deliberate and automated. It becomes fragile when the hub accumulates dependencies faster than the team can reason about them. The design goal is not maximum centralization; it is a topology whose paths and failure behavior remain understandable as the estate grows.
Private connectivity makes the topology even more subtle. A workload might use private endpoints for platform services while still relying on central DNS, route tables, and security inspection. The data path for a storage account or database can therefore depend on network controls that the application team does not directly manage. If those dependencies are not documented, application incidents can bounce between platform teams.
Service endpoints, private endpoints, direct internet paths, and centralized egress should not be mixed casually. Each pattern changes name resolution, source addresses, inspection visibility, and failure behavior. Standardizing the common patterns reduces troubleshooting time because engineers know which path should exist before they examine packets.
Telemetry is another scale boundary. Network Watcher data, firewall logs, flow logs, DNS logs, and gateway metrics can become expensive and noisy across hundreds of spokes. Central collection needs retention and query strategies that preserve evidence without making routine diagnostics impractical. The platform should provide enough visibility for workload teams to troubleshoot their own boundaries while protecting sensitive central configuration.
Change sequencing also matters. A routing policy, firewall rule, DNS update, and peering change may all be individually correct but unsafe when applied in the wrong order. Automation should model dependencies and validation between steps, especially for changes that affect shared hubs. Canary spokes or staged rollout can reduce the blast radius of platform mistakes.
When the environment is large, the healthiest sign is that common network changes are boring. New spokes receive predictable address space, peering, routes, DNS, security controls, and monitoring through repeatable automation. The hub remains important, but its behavior is well understood enough that teams can reason about failure without reverse-engineering the platform during an incident.
Platform teams should also define how exceptions expire. Temporary route bypasses, direct peerings, and firewall exclusions can become permanent because the incident that created them is forgotten. Recording an owner, reason, review date, and intended removal path keeps emergency network changes from quietly redefining the architecture.
That discipline keeps the topology understandable after years of urgent change.