Category Archives: Cloud Computing
Savings Plans and EC2 Reserved Instances both exchange flexibility for lower compute pricing, but they bind that commitment in different ways. Savings Plans commit to a consistent amount of eligible usage measured in dollars per hour. Reserved Instances commit more directly to an EC2 configuration and, depending on scope, can also provide capacity reservation. The […]
AWS Step Functions error handling is most effective when a workflow distinguishes transient failure from permanent failure. A network timeout, a throttled API, invalid business input, an authorization error, and an oversized state payload should not all receive the same retry policy. Step Functions provides retriers, catchers, explicit error names, backoff, jitter, and redrive capabilities […]
AWS Transit Gateway route tables can be used as routing domains that segment groups of VPC, VPN, Direct Connect, peering, and other supported attachments. The useful mental model is similar to virtual routing and forwarding: an attachment sends traffic into one associated transit gateway route table, and that table determines which destination attachment receives the […]
“Vertex AI Model Armor” is a useful search term, but the current Google Cloud product is Model Armor, a separate security and safety service that can protect generative-AI traffic, including Gemini requests on Google’s current Gemini Enterprise Agent Platform. Google renamed much of the Vertex AI generative-AI surface in April 2026, so architects now need […]
Vertex AI Model Garden has long been Google Cloud’s catalog for discovering and using Google, partner, and open models. In April 2026 Google renamed the wider generative-AI platform to Gemini Enterprise Agent Platform, and the current name is simply Model Garden on that platform. The older “Vertex AI Model Garden” term remains important because documentation, […]
Vertex AI Feature Store’s current architecture is built around feature data in BigQuery. Instead of maintaining a separate offline store inside Vertex AI, teams keep recent and historical feature values in BigQuery tables or views, optionally register them in the Feature Registry, and configure online stores plus feature views for low-latency serving. Current Google documentation […]
Grounding Gemini with enterprise data connects model generation to documents, websites, databases, or retrieval systems the organization controls. Current Google Cloud guidance presents several managed paths: Agent Search for Google-managed enterprise retrieval, Vertex AI RAG Engine for configurable RAG orchestration, Elasticsearch integration, external search APIs, and specialized grounded-generation APIs. These options all aim to reduce […]
Grounding Gemini with Google Search lets supported Gemini models query publicly available web information and generate responses tied to retrieved search results. Current Google Cloud documentation positions it as the default choice when an application needs current world knowledge, broad topical coverage, or up-to-date facts that are not present in the model’s static training data. […]
Vertex AI batch prediction runs large groups of model requests asynchronously instead of serving each one through a real-time endpoint. For Gemini, current Google Cloud batch inference supports input from Cloud Storage JSONL or BigQuery and can write results back to Cloud Storage or BigQuery. Batch inference is designed for high-volume, non-urgent workloads such as […]
Vertex AI context caching reduces cost and latency when Gemini requests repeatedly include the same large context. Current Google Cloud documentation distinguishes implicit caching, which happens automatically for supported models when requests share a reusable prefix, from explicit caching, where the application creates a cached-content resource and references it in later requests. The capability is […]
Vertex AI endpoint autoscaling changes the number of inference nodes behind a deployed custom or AutoML model as request demand changes. A DeployedModel defines dedicated resources such as machine type, optional accelerator, minimum replicas, and maximum replicas. Standard autoscaling keeps at least one inference node; current Google documentation also offers a preview Scale To Zero […]
Vertex AI Experiments is Google Cloud’s experiment-tracking layer for machine-learning development. Current documentation is increasingly presented under Gemini Enterprise Agent Platform Experiments, but the underlying concepts remain familiar: an experiment groups experiment runs and pipeline runs, and each run can record parameters, summary metrics, time-series metrics, artifacts, and lineage. The service is built on Vertex […]
Amazon Application Recovery Controller (ARC) routing control provides a highly available control plane for switching DNS traffic between application replicas, typically across AWS Regions. Routing controls are simple on/off states connected to specialized Route 53 health checks; changing a routing-control state changes the health status Route 53 sees and therefore which failover/weighted records receive client […]
Amazon Route 53 Resolver endpoints connect the Amazon VPC DNS resolver with DNS servers outside the VPC resolver boundary. Inbound endpoints accept DNS queries from on-premises or connected networks into VPC Resolver. Outbound endpoints send selected VPC-originated DNS queries to DNS resolvers in on-premises networks or other reachable environments according to forwarding or delegation rules. […]
Amazon S3 Multi-Region Access Points (MRAPs) provide a single global S3 endpoint in front of buckets located in multiple AWS Regions. Requests to that endpoint use the AWS global network and are routed toward an active bucket based on proximity and current routing state. For multi-Region applications, this replaces client-side region selection with one access-point […]