← All PostsAI Procurement · Vendor Due Diligence

Five Questions Before You Sign Any AI Platform Contract

Adrian DunkleySeptember 8, 20269 min read
A hand signing a document, representing the point at which an AI platform contract is committed to and the questions that should have been asked before it

Photo via Unsplash

TLDRMost AI platforms sold to boards this year are a browser interface, a set of prompts and a subscription price sitting on top of the same handful of foundation models. That means vendors are competing on demo quality, brand, RFP execution and pricing structure rather than on capability, and it means the questions that reveal a difference are not technical ones. Ask what model is underneath and what happens when it is deprecated. Ask whether the product is generative AI or machine learning, because that decides what it should be trusted with. Ask where your data goes, stays, and what audit right you hold. Ask for the error rate and the cost of an error. Ask for a customer at twelve months. Then compare vendors on the answers rather than the demos. The method has a limit, and I have set it out: it assumes you have bargaining power, and small buyers in small markets often do not.

I sat through four AI platform demos in the same month this year for three different clients. Three of the four were running Claude or GPT underneath, none of the three volunteered which, and all four demos were excellent. That is the buying environment, and it explains why the useful questions are not about capability at all.

When most products in a category are built on the same few engines, capability converges and everything else diverges. You are choosing between a vendor's understanding of their own product, their honesty about its limits, and their durability. Those are things a demo cannot show you and a question can.

What is actually underneath, and what happens when it changes

Ask the vendor to name the model and version their product runs on. Wrapping a foundation model in a good interface and a well-designed prompt structure is a legitimate product, and often a genuinely valuable one. But if that is what you are buying, you should know it, and you should price it against what it would cost to build the same thing.

The follow-up matters more than the first question. What happens when the provider deprecates that model. This is not hypothetical, it is the normal rhythm of the industry, and a vendor without an answer is telling you their product will change behaviour one day without warning. The good answer includes an evaluation framework: a set of tests they run when the base model is swapped or updated, so that a behaviour change is detected by them rather than by your client.

Generative AI or machine learning, and why the answer decides what you may trust it with

This question separates the executives who understand their purchase from the ones who do not. It gets skipped more often than the other four and it causes the expensive mistakes, so it gets more space here.

Generative AI, the family that includes ChatGPT and Claude, produces plausible text. That is the mechanism, and it is worth sitting with, because it has a direct consequence: it will always give you an answer. When the input is outside anything useful it has seen, it does not stop. It produces something that reads exactly like the answers that were correct. It also does not consult your data at all unless the vendor built retrieval into the product, which many have and some have not, and the demo looks the same either way.

Traditional machine learning works the other way round. It is fitted to your data first, and it produces a numerical prediction with confidence bounds attached. When an input falls outside its training distribution it either refuses or flags the uncertainty, because the uncertainty is part of what it computes rather than something bolted on afterwards. It does not fabricate, because generating novel content is not what it does.

The buying consequence is direct. If you are buying a tool so your team can ask questions of documents in plain language, generative AI is almost certainly right. If you are buying a tool to make classification decisions on customer data, to score credit, to flag claims, to triage applications, then being confidently wrong has a cost that lands on a real person, and machine learning is usually right. I have watched more than one Caribbean firm buy a generative product for a classification job because the demo was better, and then discover eight months later that they had no confidence bounds, no way to know when the model was out of its depth, and no defensible answer for a regulator asking how a decision was reached.

A vendor who cannot tell you which category their product belongs to has not understood their own architecture. A vendor who tells you it is both should be asked which component makes the final decision, because one of them does.

Where our data goes, and what the audit right actually is

Vendor answers on data fall on a spectrum from "your data never leaves your infrastructure" to "your data is used to train future versions of our model." Most sit in the middle: data goes to a third-party model provider, is processed there, comes back, and is logged for compliance, with the provider contractually agreeing not to train on it.

That agreement is a real protection. Its strength is entirely a function of the audit rights attached, and this is where I would push hardest. Ask what form the audit right takes, because vendors use one word for four very different things: a written attestation, a third-party certification, actual log access, or a contractual right to inspect. Those differ enormously in what they would let you discover if something went wrong, and only two of them would let you discover anything at all.

Then ask the four plain questions. Does the model train on our data. Where is it processed geographically. What is our audit right, in the specific form above. And what happens to our data if we cancel. For Caribbean firms carrying data sovereignty obligations, the geography answer is not a preference, it determines whether you can lawfully deploy the product at all, and it is worth establishing before anyone builds a business case around it.

What happens when it is wrong

Every AI system is wrong sometimes. What the vendor's answer reveals is whether they have measured that or are guessing.

Ask for the error rate on a task resembling yours, how it was measured, and under what conditions accuracy degrades. Then ask the question that follows and rarely gets asked: what does the other one percent cost, in your business, and who absorbs it. A vendor who has thought about this will answer specifically and will often volunteer the failure conditions unprompted, because they have been living with them. A vendor who has not will reach for marketing language, and marketing language will not help you the week it goes wrong.

The Five Questions, and What a Good Answer Sounds Like

  • What is underneathNames the model and version, and describes the evaluation framework that catches behaviour changes on an update
  • GenAI or MLAnswers without hesitation, and can say which component makes the final decision if both are present
  • Where does data goNames the processing geography and the specific form of the audit right, not just that one exists
  • When it is wrongGives a measured error rate on a comparable task and the conditions under which it degrades
  • Twelve-month customerProduces a named reference who will take your call, not a pilot participant or a case study

Show me somebody who has run this for a year

Ask for a real customer who has been on the product for a full twelve months and will take your call. Not a pilot participant, not a curated case study, not a logo on a slide.

A year is the interesting threshold because it covers everything a demo cannot: staff turnover, a base model update, a renewal decision made with real usage data, and the point at which the initial enthusiasm has worn off. If the vendor cannot produce one, they do not yet have a mature product, and you are the reference customer. Being the reference customer is a genuine service and should be priced as one: volume discounts, out clauses, shared IP on anything built specifically for you, and support commitments that acknowledge you will be finding their bugs.

Where this method leaves you exposed

Two limits, because this is not a complete defence.

The first is bargaining power. These questions work when a vendor needs your deal. A small buyer in a small market frequently does not have that, and a large vendor will decline to answer some of these rather than set a precedent that costs them elsewhere. The questions still earn their place in that position, because a refusal is information and belongs in the decision record, but they will not get you better terms and I would not tell a board otherwise.

The second is the twelve-month rule, which is the cleanest of the five and the one with a real cost. Applied strictly, it disqualifies every genuinely good new product, and in a field moving this fast a buyer who only ever purchases mature software will be permanently a year behind. I use it as a pricing input rather than a filter, and I have been wrong in both directions.

What the questions are really testing

None of these require you to be a technologist. They require enough hands-on time with AI tools to recognize a good answer when you hear one, which is a different and more achievable thing, and which is now part of the executive job rather than an optional interest.

Ask them, write the answers down, and compare vendors on the answers rather than the polish of the demonstration. The vendor who answers thoughtfully is the one still worth working with in eighteen months, when the market has thinned out and half the platforms in your evaluation no longer exist.

"Four demos in a month, three of them running the same model underneath, and not one of them said so. They were not hiding it. Nobody asked." - Adrian Dunkley, AI Boss

Frequently Asked Questions

How do I tell whether an AI vendor built their own model or wrapped someone else's?

Ask them to name the underlying model and version, and to say what happens when the provider deprecates it. Most platforms sold today are interfaces and prompt structures built over GPT, Claude, Gemini or Llama, which is a legitimate product but should be priced as one. A serious vendor will also have an evaluation framework that detects behaviour changes when the base model is swapped or updated, because a silent model change can alter their product's output without anyone noticing. A vendor who cannot answer either part does not fully understand what they are selling you.

What is the difference between generative AI and machine learning for a buyer?

Generative AI produces plausible text and will always produce an answer, including when it should not, and it does not consult your data unless the vendor built retrieval into the product. Traditional machine learning is fitted to your data, produces a numerical prediction with confidence bounds, and either declines or flags uncertainty when an input falls outside its training distribution. For answering natural language questions from a team, generative AI is usually right. For classification decisions on customer data, where being confidently wrong has a cost, machine learning is usually right. A vendor who cannot say which category their product falls into has not understood their own architecture.

What should I ask about where our data goes?

Four things. Does the model train on our data. Where is our data processed geographically. Do we have the right to audit that data flow, and in what form. And what happens to our data if we cancel the contract. Most vendors sit in the middle of the spectrum: data is sent to a third-party model provider, processed, returned and logged, with the provider contractually agreeing not to train on it. That agreement is a real protection, and its strength depends entirely on the audit rights attached to it.

Why do audit rights matter more than the contractual promise?

A promise not to train on your data is only as good as your ability to detect a breach of it, and data flows through a third-party model provider are not visible to you by default. Ask what form the audit right takes: a written attestation, a third-party certification, log access, or a contractual right to inspect. Those four differ enormously in what they would actually let you discover, and vendors tend to describe all of them with the same word.

What should a vendor be able to tell me about their error rate?

A specific number on a task resembling yours, how it was measured, and what the downstream cost of an error is in your business. If a vendor says the model is 99 percent accurate, the question that follows is what the other one percent costs and who absorbs it. A vendor who has measured this will have specific answers and will usually volunteer the conditions under which accuracy degrades. A vendor who has not will offer marketing language, and marketing language is not a plan for an error you will eventually have.

Why ask for a customer who has used the product for twelve months?

A pilot participant and a curated case study tell you the product demonstrates well. A customer a year in tells you whether it survived contact with real use, staff turnover, a model update and a renewal decision. If the vendor cannot produce one who will take your call, they do not yet have a mature product. That is not automatically disqualifying, but it means you are the reference customer and should price and structure the contract accordingly.

What should I negotiate if I am the vendor's first serious customer?

The things a reference customer is worth. Volume discounts against the value of the reference itself, out clauses that let you leave if the product does not mature, shared intellectual property terms on anything built specifically for you, and a commitment on support response that reflects the fact that you will be finding their bugs. Being first is a real service to a vendor and should be priced as one rather than treated as a risk you absorb for free.

Where do these five questions fail a buyer?

They assume bargaining power. A small buyer in a small market often has none, and a large vendor will simply decline to answer some of these rather than lose a deal elsewhere. In that position the questions still have value, because the refusal is itself information and should go in the decision record, but they will not get you better terms. The twelve-month question also has a real cost: applied strictly it disqualifies genuinely good new products, and a buyer who only ever buys mature software will be permanently late.

Tools for the buying conversation

AI ProcurementVendor Due DiligenceGenAI vs MLData SovereigntyAI Contracts
About the Author: Adrian Dunkley, The AI Boss

Adrian Dunkley is the founder of the Caribbean's first AI company, a distinction that placed him at the frontier of the region's technology transformation nearly two decades ago. Known across the Caribbean and internationally as the AI Boss, and recognized widely as the Godfather of Caribbean AI for the thousands of Caribbeans he has trained in artificial intelligence, he has launched and supported dozens of AI ventures spanning climate resilience, education, healthcare, agriculture, finance, and public policy. He founded StarApple AI in 2019, founded Maestro AI Labs, and chairs the Caribbean AI Risk Management Council. He holds doctorates in AI for financial inclusion and physics-based modelling and in climate physics, and sits on the buying side of AI vendor conversations across the region most weeks of the year.

Connect ↗