AI Vendor Risk Management: What to Ask Before You Sign Up 

Publication date: Jul 20, 2026

Last Published: Jul 20, 2026

Table of Contents
Read Time : 10 minutes

It takes just a few minutes to sign up for a new AI tool and, hopefully, reap its promised benefits. The problem is that AI tools don’t behave like the rest of your software. With an ordinary software-as-a-service (SaaS) product, there’s little ambiguity about what you’re sharing or where it goes. With an AI tool, what you think you’re sharing can be very different from what really gets shared and who sees it. 

This article walks through the questions that separate AI vendors that won’t jeopardize your privacy and security posture from risky ones, so you can get the productivity gains without inheriting a data problem. 

Why AI Vendor Risk Is Not the Same as Normal SaaS Risk 

Most organizations already run some version of vendor due diligence, such as whether the vendor encrypts data, how access is controlled, and what the backup story looks like. All of that still matters for AI tools, but covers less of the risk than it used to. As the National Institute of Standards and Technology (NIST) warns (PDF), generative AI brings risks that either differ from or intensify the risks of traditional software. 

With a traditional SaaS vendor, specific data goes in, sits in a locked unit, and comes back out when you need it in the exact same shape. 

With an AI tool, that certainty falls apart in two ways: 

  • The first is the data your own employees type in. When someone is racing a deadline, pasting a client list, an unsigned contract, or a chunk of source code into a chatbot is often the fastest way to get unstuck, and they rarely stop to think about where it ends up. The problem is that whoever signed your organization up for the tool has no way of knowing it happened. 
  • The second is the data the tool reaches on its own. Many AI tools earn their usefulness by connecting to the systems you already use, like your inbox, your shared drives, or your CRM. Once those connections exist, the tool can pull in far more than the few lines someone typed, and all of it travels to the vendor right along with the prompt. 

Then there’s what happens to that data once it arrives, which comes down to the vendor’s terms of service. Those terms decide how long it’s kept, who else is allowed to look, and whether your data can be used to help train a future model and effectively become part of it. When your data does become part of a model, it can theoretically later resurface in other people’s chats.  

In late 2023, researchers at Google DeepMind and several universities found that asking ChatGPT to repeat a single word over and over could make it spill back memorized pieces of its training data, including real people’s email addresses and phone numbers. This was a controlled experiment rather than something that happened to a customer in the wild, and OpenAI moved to shut the technique down, but it shows that information fed into a model can sometimes be coaxed back out, which is not a worry you have with an ordinary database.  

The numbers suggest most organizations haven’t caught up with the above-described problems. IBM’s 2025 data breach report found that 13% of organizations reported breaches of AI models or applications, that 97% of those lacked proper AI access controls, and that 63% of breached organizations either had no AI governance policy or were still writing one. In other words, organizations are buying AI faster than they’re securing it. 

If You’re a Government Contractor, the Bar Is Higher 

If you handle Controlled Unclassified Information (CUI) as a government contractor, AI vendor screening becomes a contract requirement. Under DFARS 252.204-7012, any cloud service that stores, processes, or transmits covered defense information must meet the Federal Risk and Authorization Management Program (FedRAMP) Moderate baseline or be assessed as equivalent, and that applies to AI tools like any other service. Most consumer and standard enterprise AI products don’t carry the authorization, so pasting CUI into one is a compliance incident, not a judgment call. 

Cybersecurity Maturity Model Certification (CMMC) assessments scope every system that touches CUI, so an unvetted tool also quietly widens the boundary your assessor has to examine. Two questions screen most tools quickly: 

  1. The first is whether a product is authorized for the data you handle, which is why government versions exist (Microsoft, for example, offers Microsoft 365 Copilot in its GCC High environment).  
  1. The second is whether any feature, such as web grounding or a connector, routes data outside that boundary.  

Either way, both questions are worth settling in a CMMC compliance conversation before you sign anything new, rather than discovering the problem during an assessment. 

Seven Questions to Ask Every AI Vendor 

To screen an AI vendor, you need the right questions and the discipline to insist on documented answers (a policy page, an admin guide, or a contract clause rather than a reassuring email from sales).  

If you’re still deciding whether your organization is ready for AI at all, our AI readiness checklist covers that groundwork.  

Here are the seven questions worth asking an AI vendor before you sign up: 

  1. Will you train on our data, and what is the default on our plan? This is the clearest line between consumer and business products. OpenAI’s data usage policy states that content from its consumer services may be used to train models unless you opt out, while business tiers are excluded by default. Anthropic updated its consumer terms in 2025 along similar lines. Specifically, personal Claude chats can be used for training based on a user setting, with retention up to five years for those who allow it, while business and government plans stay excluded. Defaults like these can change over time, so get the answer in writing for your plan and check it again at renewal. 
  1. Can human beings read what we submit? “Not used for training” is not the same as “never seen by a person.” Google’s privacy notice for Gemini tells consumer users plainly: “don’t enter confidential information that you wouldn’t want a reviewer to see.” Chats pulled for human review are kept for up to three years and are not deleted when you delete your activity. Ask every vendor who can read submissions, under what circumstances, and for how long. 
  1. What are your retention defaults, and which features create exceptions? Vendors increasingly advertise short or zero retention, but individual features often override the default. Google’s documentation for its enterprise agent platform supports zero data retention configurations, yet notes that the Grounding with Google Search feature stores prompts and generated output for 30 days, with no way to turn that storage off while the feature is in use. Ask specifically about memory, file uploads, connectors, and logs, and ask what happens to your data when you cancel. 
  1. What admin controls and audit logs do we get on day one? A business-ready AI product gives you single sign-on (SSO), role-based access, the ability to switch risky features off, and logs you can pull during an incident or when offboarding an employee. A recommendable vendor has all of this documented in one place you can actually read, with SSO support, retention controls, encryption, and SOC 2 audits spelled out rather than implied. If a vendor can’t show you something equivalent, assume the controls don’t exist.  
  1. Who else touches our data? An AI product is not always just one company. Behind the tool you signed up for there can be a separate model provider, a search provider, support contractors, and any connectors you turn on, each with its own data practices. Each handoff is another place your data can sit, be logged, or be exposed, and a breach at any one of them can pull in your information even though you never dealt with that company directly.  
  1. How do you defend against prompt injection and other AI-specific attacks? The Open Worldwide Application Security Project (OWASP) puts prompt injection at the top of its Top 10 security risks for large language model applications, alongside improper output handling and excessive agency, which is what happens when an AI agent holds more permissions than its task requires. A serious vendor should be able to explain in plain language how it limits what the tool can read, scopes what it can do, and tests against these attacks before release. A vendor who cannot describe any of this is telling you something, even if it is not what they meant to say. 
  1. What independent assurance can you show us today? Ask for a SOC 2 Type II report, a signed Data Processing Addendum (DPA), and ideally certification against ISO/IEC 42001, the first certifiable management standard written specifically for AI. The Cloud Security Alliance also publishes a free AI assessment questionnaire you can hand a vendor to fill out. Then watch how fast the documents arrive. A vendor with real controls produces them in days, while a vendor without them sends a glossy security overview instead. 

None of these questions require technical depth, but they do require due diligence because the honest answers usually sit three clicks deeper than the marketing page. 

Why Free AI Tools Deserve Extra Scrutiny 

Free AI tools still come at a cost, usually paid in your data, in weaker default protections, or in both. The training defaults, the human review, and the missing admin controls covered above are concentrated in free and consumer tiers, which is exactly where employees go when the company hasn’t given them an approved option. A BlackFog survey of 2,000 workers found that 49% use AI tools their employer hasn’t approved, and that most of them rely on free versions. 

This is shadow AI, and it carries a measurable price. One in five organizations in IBM’s study reported a breach caused by shadow AI, and those with high levels of it saw breach costs roughly $670,000 higher. For a company of 50 or 100 people, that is the kind of unplanned cost that can threaten the whole business. 

The fix is not a ban, which mostly pushes AI use further underground. Instead, give people an approved path good enough that no one goes looking for a free alternative, defined by an AI acceptable use policy that spells out what’s allowed and what never goes into a prompt 

Conclusion 

AI vendor risk management comes down to asking the right questions before the contract is signed, instead of after an incident report is filed. Asking them is the easy part, and a vendor worth trusting answers quickly. The harder part is reading what the answers mean and deciding which option fits your situation.  

That is where we at OSIbeyond come in. We help organizations across DC, Maryland, and Virginia put these questions to AI vendors, make sense of the answers, and match a tool’s data handling to the compliance obligations you already carry. If you’re weighing a new AI tool, or want a second look at one you already use, schedule a call with our team. 

Related Posts: