AI vendor due diligence evaluates whether a provider’s model, data practices, security, governance, operations, finances, and contract can support the risk of the intended use. A vendor that is acceptable for internal brainstorming may be unacceptable for automated eligibility decisions or access to confidential customer records. Due diligence should therefore start from the use case and consequence, not from a generic questionnaire sent to every supplier.
Within AI Governance, vendor review is a precondition for risk acceptance. The organization needs enough evidence to decide what it knows, what it cannot verify, which risks remain, and which controls belong in the contract or architecture.
NIST AI RMF provides a voluntary risk-management structure, while current OMB M-25-22 offers public procurement guidance emphasizing testing, data rights, competition, monitoring, and vendor-lock-in protections.
Start with the intended use and impact
Define users, affected people, decisions/actions, data classification, autonomy, tool permissions, deployment geography, and failure consequences.
Do not approve “Vendor X” in the abstract; approve a vendor/product/use-case combination.
A stronger review is required when the product can materially affect rights, safety, money, employment, healthcare, or privileged infrastructure.
Understand the actual model supply chain
Ask whether the vendor owns the model, resells a third-party API, fine-tunes an open model, or orchestrates several providers.
Identify infrastructure providers, subprocessors, data-labeling services, model-hosting regions, and critical open-source dependencies.
Vendor and supply-chain risk applies because the visible SaaS company may not control every critical component.
Review data use and training terms
Determine what the vendor stores, for how long, where, and whether prompts/files/outputs/feedback are used to train general models.
Ask whether zero-retention or customer-managed settings exist and whether they apply to every feature.
Review AI Training Data Rights for the vendor’s own training-data provenance and for the customer’s rights when sending data into the service.
Evaluate security beyond a certification logo
Review identity, MFA/SSO, encryption, secrets, tenant isolation, network options, logging, vulnerability management, secure development, incident response, penetration testing, access by support personnel, and model/tool-specific threats.
Ask for evidence appropriate to the risk, not necessarily every internal report.
Map vendor controls to your architecture and test the configuration you will actually deploy.
Request model and system documentation
Ask for supported use cases, limitations, evals, known failure modes, model/version lifecycle, safety behavior, fine-tuning/retrieval mechanics, and monitoring options.
For high-impact use, request enough descriptive information to complete your own risk or impact assessment.
Do not substitute a marketing model card for independent evaluation on your data and task.
Test the system with representative data before selection
Current OMB acquisition guidance encourages agencies to test proposed AI solutions and use independent evaluation data when practicable.
Private organizations should similarly run a proof-of-value against real workflows, edge cases, security/safety cases, and latency/cost constraints.
Do not let the vendor curate the only demonstration dataset used to judge fitness.
Assess change management and model lifecycle
Ask how often models/features change, whether versions can be pinned, how much notice is provided, how deprecations are handled, and whether rollback is supported.
A vendor with excellent current quality can still create governance risk if production behavior changes without traceability.
AI System Change Control should include vendor-originated changes.
Review business continuity and concentration risk
Consider provider financial health, service maturity, outage history, geographic resilience, support model, capacity commitments, rate limits, and dependency on one underlying model company.
Define what happens if the vendor exits the market, loses a key upstream provider, or discontinues your model/region.
Portability and an alternate model path can reduce concentration risk for critical systems.
Examine governance, incidents, and transparency behavior
Ask how the vendor handles safety incidents, model vulnerabilities, red teaming, abuse, privacy complaints, regulatory requests, and customer notification.
Evidence of a mature issue-management process is often more useful than a claim that the system has never had an incident.
Review who is accountable inside the vendor and how enterprise customers can escalate unresolved risk.
Turn unresolved findings into conditions, not forgotten notes
Due diligence rarely produces perfect information. Record open risks, owner, mitigation, contract clause, technical control, monitoring requirement, and review date.
If risk exceeds tolerance, use AI Risk Acceptance Decisions rather than letting procurement approval implicitly accept it.
Critical unknowns should block deployment until resolved.
AI vendor due diligence succeeds when the review remains connected to the deployed use case
The mature process understands the supply chain, tests representative behavior, verifies data/security practices, assesses change and continuity, converts requirements into contract/architecture controls, and schedules re-review when the vendor or use case changes.
Vendor diligence is not a PDF archive. It is the evidence that justifies continuing to trust an external system with a defined responsibility.
Ask for architecture diagrams that identify where customer data flows. Include model endpoints, retrieval stores, logging/telemetry, human support access, backup/DR, subprocessors, and training/feedback pipelines. A security questionnaire without a data-flow view makes it difficult to verify whether stated retention and residency promises cover every copy.
Review identity and administrator controls in the enterprise product. Require SSO/MFA, SCIM or equivalent lifecycle, role-based permissions, service accounts, audit logs, API key management, break-glass procedures, and support-access controls as appropriate. A strong underlying model does not compensate for a SaaS control plane where ex-employees retain admin access.
Assess tenant isolation through both design evidence and testing. Ask whether storage, vector indexes, fine-tuning resources, caches, and logs are tenant-scoped; whether encryption keys are shared; and how authorization is enforced. For critical systems, run tests for cross-tenant ID manipulation, object-reference guessing, and misconfigured sharing.
Evaluate AI-specific security: prompt injection, data exfiltration through tools, model extraction, unsafe code execution, malicious file handling, retrieval poisoning, and abuse of agent permissions. Ask which threats the vendor mitigates at platform level and which remain the customer’s responsibility so controls are not assumed on both sides and implemented by neither.
Check operational transparency during outages and incidents. Review status history, incident communications, service-credit process, support SLAs, and whether customers receive enough detail to assess their own downstream obligations. A vendor that provides no timely root-cause or affected-scope information can create regulatory and customer-response risk even after service is restored.
Review model evaluation evidence by cohort and task. Vendor benchmarks may not represent your language, population, document types, or safety boundaries. Ask for methodology and limitations, then run your own validation. Treat vendor metrics as input evidence, not final acceptance criteria.
Assess exit feasibility with a timed proof. Export sample data, prompts/configuration, logs, custom model artifacts, or embeddings using the vendor’s documented process. Estimate migration effort and identify proprietary formats. A vendor can claim portability while the actual export requires months of professional services.
Reassess on material events: acquisition, major funding stress, model-provider change, data-use policy change, security incident, new subprocessor, key executive/security turnover, regulatory action, or significant product architecture migration. Due diligence should have event-driven triggers in addition to an annual review calendar.
Due diligence should include accessibility and user-experience risk where the product interacts directly with people. Ask about supported languages, disability accommodations, content moderation, explanation mechanisms, and escalation to humans. A technically secure vendor can still be unsuitable if the interface excludes or misleads the population the organization must serve.
Data portability should be tested for semantic completeness, not just file export. A JSON export may omit conversation state, evaluation labels, fine-tuning relationships, tool configuration, or model-version history required to migrate successfully. Build a small exit prototype and identify which information would need manual reconstruction.
Review vendor incentives and roadmap alignment. A product optimized for broad consumer engagement may not prioritize determinism, version pinning, auditability, or regional controls important to enterprise governance. Roadmap dependency is a risk even when today’s product technically meets requirements.
Final approval should state conditions of use: allowed data classes, approved regions, prohibited decisions, required human oversight, tool limits, model versions, and monitoring obligations. This prevents a successful vendor review from being interpreted as permission to use the product for any future AI scenario.
Review the vendor’s secure-development and release process for AI-specific components. Ask how model/prompt/tool changes are tested, who can approve production updates, whether rollbacks are practiced, and how incidents feed future tests. A mature vendor should be able to explain its own change-control system rather than describing every update as continuous improvement.
Assess abuse and content-safety governance where the service is user-facing. Ask how the vendor detects harmful use, handles false positives, supports customer-specific policy, preserves privacy, and notifies customers of material safety changes. This is especially important when the organization cannot independently inspect the underlying model.
Reference checks should focus on comparable use cases. A vendor that performs well for marketing copy may have no evidence for healthcare triage or privileged coding agents. Speak with customers of similar size, geography, data sensitivity, and autonomy level where feasible, and ask about support, outages, model changes, and renewal negotiations.
Re-review should be triggered by evidence, not only the calendar, and approval conditions should remain visible to product teams.
Keep approved vendor uses scoped and visible.
Due diligence should be refreshed when the use case, data sensitivity, model capability, jurisdiction, or vendor architecture changes. A vendor that was acceptable for a low-impact pilot may need a different review once the system begins making or influencing consequential decisions.