Text analysis becomes useful when software can turn an ambiguous message into information that supports a responsible decision. A sentiment label may help route an unhappy customer, an entity extractor can identify a named product or location, and a structured model response can provide fields for downstream validation. But the appearance of a label or JSON object is not proof of correctness. A customer might express satisfaction with a delivery and frustration about a missing component in the same sentence. The application needs to preserve that distinction rather than reducing everything to a single automated judgment.
For Microsoft AI-103, text analysis is part of building dependable AI applications in Azure. The exam covers extracting entities, topics, summaries and structured outputs, analyzing sentiment and sensitive content, and adapting models for specific domains. These capabilities intersect with Foundry Tools, specialized Azure AI Language services, and generative models. The developer’s first task is to choose the operation that matches the reader’s problem, then define what happens when it produces an uncertain or incomplete result.
Start with the decision, not the model output
Consider a fictional support message: “The Contoso pump arrived in Seattle on October 8, but the replacement seal is still missing.” A support application may need to identify the organization or product, delivery city, date, problem category, customer sentiment, and whether a service agent should investigate a missing item. Each is a different assertion. The model should not infer that the entire shipment failed or that a refund was approved simply because the complaint has a negative tone.
Before implementation, define a task contract. Entity recognition should return a set of spans and classifications, with enough offsets or source context to show where the values came from. Sentiment analysis should identify the document and sentence-level opinion signals, without treating them as a diagnosis of a person’s emotional state or intent. Structured extraction should return fields that the application can type-check, and summarization should preserve the important condition that the delivery happened but the part is missing.
This separation also helps developers avoid unnecessary complexity. A deterministic parser may be better for a known invoice number or a standardized product code. A managed text-analysis service may be suitable for general named entities and sentiment. A generative model can provide flexible domain extraction or explanation, but needs additional grounding and output validation. The foundational Azure NLP decision framework explains how task selection depends on business meaning. The implementation layer turns that decision into testable software behavior.
Write down what the application will do with the result. A negative sentiment score might trigger a queue prioritization, but should not automatically deny service or label the customer as abusive. A recognized city is not proof of delivery jurisdiction. A summary cannot replace the original message when a human investigator needs the exact wording. Good contracts distinguish suggested insight from authorized business action.
Choose a stable specialized client before adding generative extraction
Azure AI Language exposes specialized capabilities for sentiment analysis, opinion mining and named-entity recognition. Microsoft documents Python examples using the azure-ai-textanalytics package, including version 5.2.0 in its current quickstarts. Those examples use TextAnalyticsClient with an Azure resource endpoint and authorized credential. Newer SDK families and Foundry integrations may offer different client classes or model paths, so this example intentionally stays within one documented interface rather than mixing snippets from incompatible packages.
A real Azure resource must be created and configured by someone authorized to operate it. For local development, the service documentation demonstrates environment variables for LANGUAGE_ENDPOINT and LANGUAGE_KEY. In production, use secure server-side credentials or supported Microsoft Entra workload identity rather than exposing keys to browsers, public repositories or client bundles. Treat the language endpoint as a capability with network, authentication and regional constraints, not a harmless public URL.
The following small Python example analyzes one synthetic support message, using sentiment with opinion mining and named-entity recognition. Install azure-ai-textanalytics==5.2.0 in an authorized environment first. This illustrates SDK method calls and error handling, not a claim that any customer data was processed or Azure service invoked while writing this article.
import os
from azure.ai.textanalytics import TextAnalyticsClient
from azure.core.credentials import AzureKeyCredential
client = TextAnalyticsClient(
endpoint=os.environ["LANGUAGE_ENDPOINT"],
credential=AzureKeyCredential(os.environ["LANGUAGE_KEY"]),
)
message = (
"The Contoso pump arrived in Seattle on October 8, "
"but the replacement seal is still missing."
)
sentiment = client.analyze_sentiment([message], show_opinion_mining=True)[0]
entities = client.recognize_entities([message])[0]
if sentiment.is_error or entities.is_error:
raise RuntimeError("Text analysis did not complete successfully")
print("Overall sentiment:", sentiment.sentiment)
print("Sentence labels:", [x.sentiment for x in sentiment.sentences])
for entity in entities.entities:
print("Entity:", entity.text, entity.category, entity.confidence_score)
The sample prints document sentiment, sentence labels and recognized entity text, category and confidence. It deliberately does not automatically write anything to a customer account. The exact number of returned entities and opinion targets can vary by supported model version and language. A production caller should also respect service quotas, request size limits, per-document errors, input language support and handling of sensitive source data. Failure of one analysis task should not silently overwrite or invalidate the result of another task.
Opinion mining can provide more granular targets and assessments than an overall sentiment label. This is useful when a review praises the product but criticizes shipping or support. Evaluate whether the returned target spans genuinely identify the subject of the opinion. A sentence about a defective seal should not be treated as a verified mechanical failure simply because the model labeled a negative assessment. The user report is evidence of a complaint, not conclusive technical diagnosis.
Interpret sentiment as a signal rather than a business verdict
Sentiment analysis commonly returns positive, neutral, or negative predictions along with confidence scores at the document and sentence levels. These scores express the service’s classification confidence, not an objective measurement of happiness, customer worth or future behavior. A short message, sarcasm, mixed praise and criticism, cultural language variations or domain-specific terminology can make interpretation more difficult. A support workflow should therefore combine the signal with actual facts from the message and known service policies.
Build a test collection that includes mixed sentiment, short ambiguous requests, polite but serious complaints, and messages with positive words used sarcastically. Inspect sentence-level differences before trusting a document-level summary. For example, “Thanks for sending the order so quickly, but the seal is missing” can contain praise and an important problem. An automated workflow that reads only the positive part may fail the customer. Human review should be available for uncertain or high-consequence routing decisions.
Opinion mining adds target and assessment relationships. In a review of a software application, one sentence may praise interface speed and criticize account recovery. If the output correctly locates the target but confuses its assessment, dashboards built on it will distort product decisions. Evaluate with labeled examples specific to the domain instead of relying only on general sentiment demo sentences.
Do not use a sentiment label as permission to take irreversible or sensitive action. The system should not infer an employee’s performance, a customer’s honesty, or a medical condition from linguistic sentiment alone. If sentiment assists with escalation, define the authorized decision separately and measure false positives and false negatives against operational outcomes.
Recognize entities without discarding their evidence
Named-entity recognition can identify people, organizations, places, dates, quantities and other classes supported by a model. The result usually contains spans or offsets that connect the proposed entity to the original input. Keep those references when software needs to verify a label or value. A downstream object such as delivery_city may be populated from a recognized location, but only after the application establishes that the location refers to delivery rather than company headquarters, a previous address, or an event venue.
General entity types do not automatically equal business-specific fields. A model may recognize a product name as an organization, or a date as a generic time reference. A domain extractor still needs context and validation. For example, the support message might mention an arrival date, an order date and a promised replacement date. Returning all three as “DateTime” entities is a useful perception result but not a final answer to “When was the replacement due?”
When matching a named entity to an internal record, use a separate authorized lookup. A detected product label should not automatically grant access to a customer record. Validate the match using the allowed business identifiers and keep the source span and confidence for review. This is also where sensitive-data handling matters: some extracted names, locations and account details require retention and role restrictions different from a generic sentiment score.
If a workflow handles specialized terms, compare general NER with a purpose-built extraction approach. A schema-driven generator might identify an item category and expected resolution, but it can fabricate a plausible missing field. Validate its output types and known values, and include an explicit unknown or not-present state instead of forcing every record to look complete.
Use structured outputs as contracts, not as proof of truth
Generative language models can extract topics, summarize messages and return JSON conforming to a requested schema. The benefit is a predictable software interface, but the schema only constrains representation. A valid JSON object can still contain a fabricated transaction ID, the wrong cause of a failure, or a sentiment label that is unsupported by the source. A dependable system validates both structure and meaning.
For our support message, a useful business object might contain product_name, delivery_location, arrival_date, missing_component, requested_action, and evidence_quote. Each should have a type and clear instructions for absent information. The system should avoid inventing requested_action if the customer merely describes a problem without explicitly asking for a refund. In that case, a null or unresolved field is more honest than a guessed instruction.
Explore the differences between schema validation, domain validation and policy enforcement. Schema validation checks that an ISO date string is present when expected. Domain validation checks whether the date makes sense and whether named values are permitted. Policy enforcement checks what the business is allowed to do with those values. Azure OpenAI structured outputs can help constrain response formats, but a field that passes the schema still requires evidence and authorization for consequential uses.
Build negative tests. Supply a message with no order identifier, with multiple dates, with two products, and with a negated instruction such as “Do not cancel the replacement.” The correct parser must not discard that negation. Inspect how the model handles adversarial or irrelevant text embedded in the message and whether it inappropriately treats customer content as higher-priority system instructions. Keep the source and parsed result connected for troubleshooting.
Summarize domain-specific text without erasing qualifications
Summarization is particularly risky when a shorter description becomes the version of events shown to an employee. A summary of a support exchange can omit that a customer already completed a required step or that an agent promised only a conditional exception. In compliance and operations settings, these omissions may change how the case is handled. Define what the summary is for before setting the length or style.
For case routing, a short summary may need the current problem, relevant order state, prior troubleshooting and next authorized action. For compliance review, it may need source references, dates, exclusions and unresolved questions. A model instructed only to “make it concise” may remove the very qualifications a reviewer needs. Ask it to preserve decision-relevant facts, uncertainties and contradictory evidence; then verify those attributes against the source document or conversation.
Use a controlled evaluation set with examples where a seemingly small omission alters meaning. Include “refund approved if the item is returned” versus “refund approved,” “service may be unavailable” versus “service is unavailable,” and similar distinctions. Evaluate factual consistency, coverage of required issues, and absence of invented actions independently of readability. An attractive summary can be worse than a slightly longer accurate one.
If the application uses retrieval-augmented context, a summary must retain the permitted source boundary. A user cannot gain access to a private record simply because it was included in a model’s context window. Summaries should not merge confidential records with public support notes or reveal restricted account data to unrelated users.
Protect sensitive content and separate safety from tone
A frustrated user is not necessarily using disallowed language, and a polite request can still contain an unsafe instruction or a sensitive secret. Sentiment, toxicity or content safety, personally identifiable information detection, and policy compliance are separate classifiers or checks. Treat them as different controls. Use the current supported services for the intended purpose and evaluate which false positives could prevent a legitimate user from receiving support.
Minimize unnecessary submission of personal data. Redact or replace account identifiers where the task does not require them, retain the original securely if business records demand it, and restrict access to raw content and derived annotations. An extracted entity can itself reveal sensitive information, so do not assume that a JSON array of detected names is safe to log. Prefer structured diagnostic information and correlation IDs over copying entire customer conversations into general telemetry.
Model prompt injection is also a concern. An incoming customer message might include instructions for an AI system to ignore its rules, expose an internal list, or approve a transaction. That text is customer data, not trusted application policy. Validate every proposed downstream action independently. If a classifier or model indicates that a request might be unsafe, the application still needs a defined rejection or human-review path appropriate to the real business workflow.
Where supported, use secure identity and role-based access for service connections. The design of managed identity for AI applications helps reduce stored credentials in suitable workloads. The calling identity’s permission to invoke a model does not authorize the application to expose customer records or make account changes. Separate service permissions from user permissions and keep audit evidence proportional to risk.
Evaluate and operate a full text-analysis workflow
An initial test can run the documented sentiment and NER methods on invented messages. A fuller test should include both correct and incorrect model outcomes, a range of input languages, ambiguous names, mixed sentiment, multiple possible dates, missing business fields, long messages and input that should be rejected or sent to review. For each input, record the expected entity spans or values and the important conditions that should survive interpretation.
Measure at least three failure types: wrong perception of the text, wrong extraction into a business field, and wrong business action following a seemingly valid extraction. These failures need different fixes. A missed named entity may require a better model or source normalization; a fabricated JSON value requires grounding or validation; an unauthorized transaction requires application enforcement even if the model’s interpretation is otherwise correct.
Instrument the flow so that an authorized operator can trace an output to input version, model/service configuration, validation result and policy decision without retaining more private text than necessary. Handle rate limits, service errors and unsupported languages explicitly. Avoid retrying an invalid input indefinitely, and ensure a rejected or timed-out text-analysis request does not silently turn into a false “neutral” sentiment or empty “no entities” result.
For AI-103 preparation, the most important exercise is to explain why each operation exists. Sentiment can help with review and routing, entities can inform a structured lookup, and generative output can summarize or extract flexible fields. None should be allowed to decide sensitive outcomes without the surrounding application logic. A successful text-analysis system leaves its evidence inspectable, preserves uncertainty and fails safely when meaning cannot be established.