Category Archives: Cloud Computing
Azure Monitor Data Collection Rules, commonly called DCRs, define how supported telemetry enters Azure Monitor: what data sources are collected, which streams they produce, which transformations are applied, and where the results are sent. They replace many one-off collection configurations with a centralized resource that can be versioned, associated with workloads, and managed through automation. […]
Azure user-defined routes, or UDRs, let administrators override parts of the platform’s default routing behavior. They are central to hub-and-spoke networks, forced tunneling, network virtual appliances, Azure Firewall architectures, and hybrid connectivity. A small route-table change can therefore affect a large number of workloads even when every VM, firewall rule, and application process remains healthy. […]
Cloud Run services and Cloud Run jobs use the same container-oriented execution environment, but they solve different control-flow problems. A service waits for requests or events and scales instances around demand. A job starts tasks that run to completion and then stops. Choosing between them is less about which product is newer and more about […]
The useful question is not whether a small language model or a large language model is “better.” It is which model class satisfies the quality, latency, privacy, tool-use, context, and cost requirements of a particular workload. Modern model catalogs now span compact models that can run close to the user and much larger models designed […]
Structured output turns a model response from “text that usually looks right” into data that downstream software can treat as an explicit contract. That distinction is especially important for agents. An agent does not merely display prose; it may select tools, create records, route work, trigger approvals, or pass intermediate results to another model. If […]
AI token cost forecasting is less about predicting one invoice and more about understanding how workload behavior creates spend. Token-based services usually charge separately for input and output, may price cached context differently, and can vary substantially by model tier. An agent adds another source of variation: one user request can trigger retrieval, several model […]
Tool calling makes an AI agent useful because the model can reach systems that contain current data or perform actions. It also creates a distributed-systems problem. Tools time out, credentials expire, APIs rate-limit requests, schemas change, responses are incomplete, and side effects may succeed even when the caller never receives confirmation. An agent that treats […]
Human-in-the-loop AI is often described as adding an approval button to an automated workflow. That is too narrow. A meaningful human control defines when an AI system must stop, what evidence a reviewer sees, what authority the reviewer has, what happens after approval or rejection, and how the decision is recorded. Without those details, “human […]
First-stage retrieval is designed to search a large corpus quickly. It may use vector similarity, keyword search, hybrid retrieval, or another indexing technique to return a candidate set. Reranking is a second stage that examines those candidates more carefully and reorders them according to their relevance to the specific query. The distinction matters because the […]
AI red-team test cases are structured adversarial scenarios designed to reveal how a generative AI system behaves when users, retrieved content, or tools push it toward unsafe, unauthorized, deceptive, or unreliable behavior. Effective red teaming is not a collection of provocative prompts. It is a repeatable engineering practice that connects a threat model to test […]
Chunking by document structure divides source material according to headings, paragraphs, lists, tables, and other semantic boundaries instead of cutting text at fixed character or token intervals. The goal is to preserve meaning so each retrieved chunk carries enough local context to answer a question without dragging unrelated sections into the model. In Agentic AI […]
A context window is not an unlimited notebook attached to a model. It is a bounded working set that must hold instructions, conversation state, retrieved evidence, tool descriptions, tool results, examples, and the user’s current request at the same time. When teams treat that space as free, the predictable symptoms are higher latency and cost, […]
Model monitoring is easy to reduce to a chart labeled “drift,” but production monitoring is really a decision system. A team chooses a baseline, chooses the production data to compare against it, selects a distance metric, defines thresholds, and then decides what an alert should cause humans or automation to do. Google’s current Model Monitoring […]
Safety filters are one of the most misunderstood parts of a generative-AI architecture because teams often treat them as a universal “safe mode.” In practice, the Gemini API exposes configurable content-safety controls for defined harm categories, while Google also applies non-configurable protections for certain prohibited content. The application chooses thresholds for supported configurable categories and […]
“Vertex AI Search grounding” now spans two generations of Google Cloud naming. In April 2026 Google renamed Vertex AI Search to Agent Search on Gemini Enterprise Agent Platform. The console and some APIs still expose older names, and the retrieval object used in Gemini requests can still be labeled VertexAISearch. That makes current architecture work […]