Most of your AI risk now comes from tools you buy, not models you build. An AI vendor is not just a supplier, they are handling your data, shaping decisions, and inheriting a slice of your risk and your compliance burden. Evaluating one well is the difference between buying capability and buying a liability. The good news: it is a repeatable process, not a leap of faith, if you know what to look at.
People often evaluate AI tools on the demo: does it look impressive, does it do the thing. That is the least important part. The demo shows you the upside on a good day. What you actually need to know is what happens on a bad one, and what you are signing up to. Here is how to do that.
Why this matters more for AI
Buying AI is not like buying ordinary software:
- The tool often processes your data in ways that are hard to see, sometimes to train models you cannot audit.
- The EU AI Act applies to AI you use, not just AI you build, so a bought tool can pull you into deployer obligations.
- AI features arrive quietly inside tools you already own, which is how shadow AI spreads. A bought AI tool that nobody evaluated is exactly the gap governance is meant to close.
So vendor evaluation is not procurement box-ticking. It is where a lot of your AI exposure is decided.
What to actually evaluate
Score every candidate against the same criteria. These are the ones that earn their place:
| Area | What to check | Why it matters |
|---|---|---|
| Your data | Do they train on it? Retention, location, sub-processors, export | This is where the biggest, least visible risk lives |
| Security | SOC 2 / ISO 27001, encryption, pen tests, DPA | You inherit their weaknesses |
| Compliance | GDPR, and increasingly ISO 42001 for AI specifically | Their posture becomes part of yours |
| Model transparency | Which model, versioning, can it change under you, accuracy claims | ”It’s AI” is not an answer; you need specifics |
| Oversight & control | Human-in-the-loop options, audit logs, permissions | Required for anything sensitive or high-risk |
| Reliability | Uptime SLA, latency, rate limits, real accuracy on your data | Demos hide the bad days |
| Lock-in | Open standards, data and prompt export, portability | How hard is it to leave if it goes wrong |
| Viability | Funding, track record, support, references | Will they still be here, and answer the phone |
| Cost model | Usage pricing, overage surprises, price-change terms | AI bills scale in ways that surprise people |
You will not weigh these equally for every purchase. A low-stakes internal assistant needs less scrutiny than a tool that touches customer data or drives a decision. Match the depth to the risk tier.
The data questions that matter most
If you only press hard on one area, make it data. Ask, and get the answers in the contract, not the sales deck:
- Do you train your models on our data, by default or ever?
- Where is our data stored and processed, and who are your sub-processors?
- How long is it retained, and how do we delete it?
- How do we export everything and leave?
A credible vendor answers these crisply and puts them in writing. Vague, shifting, or “trust us” answers are the single biggest red flag in AI procurement.
Run a proof of concept on your own data
The demo is their best case. A proof of concept on your data, your messy inputs, your edge cases, is the only way to see real accuracy and fit. Define what “good enough” looks like before you start, so you are measuring against a bar rather than a vibe. This is the same discipline as building an evaluation harness before you ship, applied to a purchase.
Red flags
- Won’t put data terms in writing. “We don’t train on your data” means nothing if it is not contractual.
- No security evidence. No SOC 2, ISO 27001 or DPA to show.
- Black-box accuracy claims. Big numbers with no methodology, and no way to test on your data.
- No export path. If you cannot get your data and workflows out, you are locked in by design.
- Hand-waving on the EU AI Act. If a tool is high-risk in your use and the vendor cannot speak to logging, documentation or oversight, that burden lands on you.
Where this fits in governance
Evaluation is not a one-off gate at purchase. A bought AI tool belongs in your AI system register, gets risk-classified like anything else, and its vendor becomes part of the supplier management inside your AI Management System. The evaluation is how it enters that system with its risks understood, rather than slipping in unseen. This is the buy side of the coin; the mirror image, proving the trustworthiness of the AI you sell, is AI assurance.
The short version
Evaluating an AI vendor is mostly about looking past the demo at what you are actually signing up to: what happens to your data, how it is secured, whether it helps or hurts your own compliance, and how hard it is to leave. Score every candidate on the same criteria, press hardest on data and get the answers in writing, and prove it on your own data before you commit. Do that and you buy capability, not a liability.
Weighing up an AI tool and want a second pair of eyes, or help building the governance around the ones you keep? Talk to us.
For the full picture, see the AI governance guide.
Frequently asked questions
What should you look for when evaluating an AI vendor?
The essentials are: what happens to your data (training, retention, location, sub-processors), the vendor's security and compliance posture (SOC 2, ISO 27001, ISO 42001, GDPR), how transparent they are about the model and its accuracy, whether they support human oversight and audit logging, how easily you can export your data and avoid lock-in, and their reliability and long-term viability. Score every candidate against the same criteria, and always run a proof of concept on your own data.
What should I ask an AI vendor about my data?
Ask whether they train their models on your data by default, and get the answer in the contract, not just marketing copy. Then ask where data is stored and processed, how long it is retained, who their sub-processors are, whether data is encrypted in transit and at rest, and how you get your data out if you leave. Vague or shifting answers on any of these are a red flag.
Does the EU AI Act apply to AI tools I buy?
Yes. The EU AI Act applies to organisations that use or deploy AI, not just those that build it. If a tool you buy is used in a high-risk context, deployer obligations attach to you regardless of who made it. So part of evaluating a vendor is checking whether they help you meet those obligations, with things like logging, documentation and human-oversight controls, or leave you to carry the burden alone.
How do you avoid vendor lock-in with AI tools?
Favour vendors that use open standards (such as the Model Context Protocol for tool and data access), let you export your data and prompts in a usable format, and do not bury your workflows in proprietary formats you cannot leave. Before committing, ask concretely how you would migrate off, and make sure the answer is not "you cannot".
How do you check an AI vendor is secure and compliant?
Ask for evidence, not assurances: current SOC 2 Type II or ISO 27001 certification, a data processing agreement, penetration-test summaries, and, increasingly, ISO 42001 for AI management specifically. Confirm GDPR/UK GDPR compliance and where data is processed. A credible vendor produces this readily; one that cannot is telling you something.
Building something you need to govern?
Start with a fixed-scope AI Opportunity & Risk Audit.
Meet an Expert