Bedrock prompt routers are named resources used by Amazon Bedrock intelligent prompt routing to choose among supported models. A router encapsulates candidate models, a fallback model, and routing criteria; applications invoke the router rather than implementing their own request classifier. Current AWS APIs include CreatePromptRouter, GetPromptRouter, ListPromptRouters, and DeletePromptRouter alongside AWS-provided default routers.
Within Generative AI on AWS, the router resource is the production configuration boundary around Bedrock Intelligent Prompt Routing. Intelligent routing describes the behavior; the prompt router is the versioned/configured object applications actually govern.
Routers should be treated like load-balancer or API-gateway policy: named, tagged, evaluated, deployed through infrastructure code, and changed deliberately.
A configured router defines candidates and a fallback
The CreatePromptRouter API accepts candidate model ARNs, a fallback model ARN, a router name, and a routing criterion.
The fallback model is the higher-quality/default path when the configured response-quality threshold is not satisfied by the other candidate.
Document why each model belongs in the router and what quality/cost role it serves instead of adding models simply because the API allows them.
responseQualityDifference is a policy knob, not a magic accuracy score
The routing criterion expresses the tolerated predicted quality difference used by the router.
A lower or higher setting changes the share of traffic that can be sent to the lower-cost candidate versus the fallback.
Calibrate the value through application evaluation and record the resulting model-selection distribution, quality, and cost before promoting it.
Router names should represent workload intent
Use names such as support-prod-router or doc-summarization-prod, not names that hard-code one candidate model.
The logical purpose should survive model changes, while tags/description record owner, environment, cost center, and evaluation version.
This makes resource inventories and IAM policies understandable even after the underlying model pair changes.
Router changes should use replacement/versioning discipline
A production application should not experience an untested threshold/model change merely because an engineer edited the router configuration.
Create/evaluate a candidate router, route canary traffic, compare metrics, then switch application configuration to the approved ARN where the service workflow requires replacement.
Keep the previous router available long enough for rollback if model availability and cost allow.
Default and custom routers need separate governance
Default routers are AWS-managed resources and can evolve as AWS adds supported model options.
Configured routers give the customer more explicit control over candidate/fallback choices and criteria.
Applications that require reproducible behavior should document whether they use a default or custom router and which router ARN/version-like configuration was active for each evaluation.
IAM should constrain who can create or delete routers
A prompt-router change can shift substantial spend and output behavior across models.
Limit control-plane permissions to the platform/release pipeline while application identities receive only the invocation permissions they need.
Use CloudTrail and tagging to audit who created/deleted a router and which application owns it.
Runtime metrics should be attributed to router and selected model
Observability needs both the logical router and the actual model chosen for each request.
Track total requests, model-selection rate, task success, latency, tokens, cost, guardrail interventions, and error rate.
A router can look cheaper on average while sending the hardest prompts to a model that is temporarily overloaded or producing more retries.
Fallback behavior should be tested at the decision boundary
Build prompts near the quality threshold and see whether small wording changes cause model switching.
Measure whether that variability matters to users, tool-call behavior, structured output, or deterministic downstream workflows.
If a workflow requires strict consistency, a fixed model may be preferable even when intelligent routing saves money elsewhere.
Routers should be part of cost allocation
Tag custom routers and correlate invocation records with workload/cost center.
If several applications share one router, internal chargeback becomes less precise and teams can affect each other’s cost/quality policy.
Prefer separate application routers when ownership, evaluation, or budget differs materially.
Router lifecycle should follow model lifecycle
When a candidate model is retired, unavailable in a Region, or superseded, the router needs evaluation/update.
Monitor supported-model documentation and create a migration backlog before retirement dates affect production.
The router abstraction reduces application code changes, but it does not eliminate responsibility to keep the model set current.
Prompt routers succeed when dynamic selection is operated like production routing policy
The mature platform versions candidate/fallback choices, calibrates criteria, secures control-plane changes, observes selected models, tags cost ownership, canaries updates, and keeps fixed-model exceptions for workflows that require determinism.
A router should centralize model-selection policy without hiding the consequences of that policy from application owners.
Router inventory should include the exact candidate model ARNs and fallback model rather than relying on human-readable console labels. Model version changes can be subtle, and two routers with similar names may behave differently because one points at an older release. Export router configuration into source control or inventory so drift is visible.
Evaluation datasets should be attached to the router lifecycle. A routing criterion is meaningful only relative to the prompts used to calibrate it. Keep a fixed regression set for the workload plus a rolling production-failure set, and rerun both before changing candidates or thresholds.
Router rollback should be one configuration change away. Applications should reference the router ARN through environment/configuration or a platform abstraction, not compile it into many services. This makes it possible to switch back to a known-good router quickly if a new candidate produces poor behavior.
Default router adoption should be monitored for model changes. AWS can evolve default offerings as model families change. If strict reproducibility matters, regularly snapshot which model options the default router can select and consider moving to a configured router once the workload becomes production-critical.
Router-based cost optimization should account for token-output differences. The cheaper model may answer with more tokens or trigger additional tool calls, which can reduce savings. Calculate end-to-end workflow cost by selected model rather than simply multiplying request count by list price.
Prompt-router metrics should be segmented by prompt class, customer tier, language, and workflow. A router can perform well on general support questions but poorly on code or regulated domain prompts. Segmenting results reveals where fixed-model exceptions or separate routers are justified.
Security policy should apply outside the router. IAM decides who can invoke the resource, Bedrock Guardrails can evaluate content, and application authorization decides which tools or data may be used. The router should not be asked to distinguish “safe customer” from “unsafe customer” or “authorized operation” from “unauthorized operation.”
Infrastructure changes should include deletion protection at the process level. A DeletePromptRouter call can break any application that references that ARN. Use IaC review, dependency inventory, and staged decommissioning so an unused-looking router is not removed while a batch job or secondary environment still depends on it.
Prompt routers are most effective when one workload has a clear quality/cost trade-off. If an application contains several radically different tasks, separate routers can be easier to evaluate than one router receiving everything. Narrow workload scope makes the routing criterion easier to reason about and monitor.
Configuration review should capture the expected traffic split under the evaluation dataset. After deployment, compare actual model-selection share with that expectation. A large drift can indicate the production prompt mix changed, which may warrant recalibration even when the router resource itself did not change.
Keep router dependencies in inventory. Prompt-management resources, guardrails, inference profiles, application gateways, and observability can all reference or assume a router. Decommissioning should confirm every consumer has moved and that dashboards/budgets no longer depend on the old resource before deletion.
For high-stakes workflows, define explicit non-routed paths. A payment approval, legal determination, or irreversible automation step may require one tested model even if surrounding summarization can use the router. Router adoption should be selective by task, not mandatory by platform policy.
Router changes should be peer-reviewed with both technical and product metrics. The same threshold that looks acceptable to a platform engineer may alter escalation rate or tone in ways a product owner notices immediately. Treat routing policy as a user-experience control as well as a cost-control resource.
For canary rollout, route a small percentage of eligible workload to the candidate router while a control cohort stays on the previous router or fixed model. Compare quality, tool correctness, latency, and total cost before moving the rest of traffic.
Prompt-router governance should also document unsupported workloads. If multilingual, safety-critical, highly specialized, or deterministic tasks are excluded, encode those exclusions in application routing logic so they do not accidentally enter the dynamic router later.
Router policy should also have a clear retirement process. When an application no longer needs dynamic selection, remove traffic first, validate that no batch or secondary environment still references the ARN, archive the evaluation/configuration evidence, then delete the router. Resource cleanup should preserve enough history to explain past production behavior.
The platform should publish which prompt routers are approved for which workload classes, models, and Regions. This avoids teams discovering a router in the console and reusing it for a domain whose quality or compliance characteristics were never evaluated.
Router changes should be evaluated against representative prompts rather than only aggregate latency. A lower-cost or faster model is not an improvement if routing shifts difficult requests to a model that degrades accuracy, safety, tool use, or structured-output reliability.