Photo via Unsplash
A wrapper is a platform built as an interface and a prompt structure over someone else's foundation model, usually GPT, Claude, Gemini or Llama. That is a legitimate product. It should be priced as one, and you should know when you are buying it.
Generative AI versus machine learning is the distinction that decides what a tool may be trusted with. GenAI produces plausible text and will always produce an answer, including when it should not, and it does not consult your data unless retrieval was built in. Traditional ML is fitted to your data, returns a prediction with confidence bounds, and declines or flags uncertainty when an input falls outside its training distribution. Buy GenAI for answering questions from documents. Buy ML for classification decisions about customers, where being confidently wrong lands on a real person.
An audit right is one of four very different things vendors describe with the same word: a written attestation, a third-party certification, actual log access, or a contractual right to inspect. Only two of them would let you discover anything.
The five questions. What model is underneath and what happens when it is deprecated. GenAI or ML. Where does our data go, stay, and what is the audit right. What is the measured error rate and what does an error cost us. Who has run this for twelve months and will take my call.
The limit. All five assume bargaining power. A small buyer in a small market often has none, and a large vendor will decline rather than set a precedent. The refusal is still information, so record it.
I sat through four AI platform demos in the same month this year for three different clients. Three of the four were running Claude or GPT underneath, none of the three volunteered which, and all four demos were excellent. That is the buying environment, and it explains why the useful questions are not about capability at all.
When most products in a category are built on the same few engines, capability converges and everything else diverges. You are choosing between a vendor's understanding of their own product, their honesty about its limits, and their durability. Those are things a demo cannot show you and a question can.
What is actually underneath, and what happens when it changes
Ask the vendor to name the model and version their product runs on. Wrapping a foundation model in a good interface and a well-designed prompt structure is a legitimate product, and often a genuinely valuable one. But if that is what you are buying, you should know it, and you should price it against what it would cost to build the same thing.
The follow-up matters more than the first question. What happens when the provider deprecates that model. This is not hypothetical, it is the normal rhythm of the industry, and a vendor without an answer is telling you their product will change behaviour one day without warning. The good answer includes an evaluation framework: a set of tests they run when the base model is swapped or updated, so that a behaviour change is detected by them rather than by your client.
Generative AI or machine learning, and why the answer decides what you may trust it with
This question separates the executives who understand their purchase from the ones who do not. It gets skipped more often than the other four and it causes the expensive mistakes, so it gets more space here.
Generative AI, the family that includes ChatGPT and Claude, produces plausible text. That is the mechanism, and it is worth sitting with, because it has a direct consequence: it will always give you an answer. When the input is outside anything useful it has seen, it does not stop. It produces something that reads exactly like the answers that were correct. It also does not consult your data at all unless the vendor built retrieval into the product, which many have and some have not, and the demo looks the same either way.
Traditional machine learning works the other way round. It is fitted to your data first, and it produces a numerical prediction with confidence bounds attached. When an input falls outside its training distribution it either refuses or flags the uncertainty, because the uncertainty is part of what it computes rather than something bolted on afterwards. It does not fabricate, because generating novel content is not what it does.
The buying consequence is direct. If you are buying a tool so your team can ask questions of documents in plain language, generative AI is almost certainly right. If you are buying a tool to make classification decisions on customer data, to score credit, to flag claims, to triage applications, then being confidently wrong has a cost that lands on a real person, and machine learning is usually right. More than one Caribbean firm has bought a generative product for a classification job because the demo was better, then discovered eight months later that it had no confidence bounds, no way to know when the model was out of its depth, and no defensible answer for a regulator asking how a decision was reached.
A vendor who cannot tell you which category their product belongs to has not understood their own architecture. A vendor who tells you it is both should be asked which component makes the final decision, because one of them does.
Where our data goes, and what the audit right actually is
Vendor answers on data fall on a spectrum from "your data never leaves your infrastructure" to "your data is used to train future versions of our model." Most sit in the middle: data goes to a third-party model provider, is processed there, comes back, and is logged for compliance, with the provider contractually agreeing not to train on it.
That agreement is a real protection. Its strength is entirely a function of the audit rights attached, and this is where I would push hardest. Ask what form the audit right takes, because vendors use one word for four very different things: a written attestation, a third-party certification, actual log access, or a contractual right to inspect. Those differ enormously in what they would let you discover if something went wrong, and only two of them would let you discover anything at all.
Then ask the four plain questions. Does the model train on our data. Where is it processed geographically. What is our audit right, in the specific form above. And what happens to our data if we cancel. For Caribbean firms carrying data sovereignty obligations, the geography answer is not a preference, it determines whether you can lawfully deploy the product at all, and it is worth establishing before anyone builds a business case around it.
What happens when it is wrong
Every AI system is wrong sometimes. What the vendor's answer reveals is whether they have measured that or are guessing.
Ask for the error rate on a task resembling yours, how it was measured, and under what conditions accuracy degrades. Then ask the question that follows and rarely gets asked: what does the other one percent cost, in your business, and who absorbs it. A vendor who has thought about this will answer specifically and will often volunteer the failure conditions unprompted, because they have been living with them. A vendor who has not will reach for marketing language, and marketing language will not help you the week it goes wrong.
The Five Questions, and What a Good Answer Sounds Like
- What is underneathNames the model and version, and describes the evaluation framework that catches behaviour changes on an update
- GenAI or MLAnswers without hesitation, and can say which component makes the final decision if both are present
- Where does data goNames the processing geography and the specific form of the audit right, not just that one exists
- When it is wrongGives a measured error rate on a comparable task and the conditions under which it degrades
- Twelve-month customerProduces a named reference who will take your call, not a pilot participant or a case study
Show me somebody who has run this for a year
Ask for a real customer who has been on the product for a full twelve months and will take your call. Not a pilot participant, not a curated case study, not a logo on a slide.
A year is the interesting threshold because it covers everything a demo cannot: staff turnover, a base model update, a renewal decision made with real usage data, and the point at which the initial enthusiasm has worn off. If the vendor cannot produce one, they do not yet have a mature product, and you are the reference customer. Being the reference customer is a genuine service and should be priced as one: volume discounts, out clauses, shared IP on anything built specifically for you, and support commitments that acknowledge you will be finding their bugs.
Where this method leaves you exposed
Two limits, because this is not a complete defence.
The first is bargaining power. These questions work when a vendor needs your deal. A small buyer in a small market frequently does not have that, and a large vendor will decline to answer some of these rather than set a precedent that costs them elsewhere. The questions still earn their place in that position, because a refusal is information and belongs in the decision record, but they will not get you better terms and I would not tell a board otherwise.
The second is the twelve-month rule, which is the cleanest of the five and the one with a real cost. Applied strictly, it disqualifies every genuinely good new product, and in a field moving this fast a buyer who only ever purchases mature software will be permanently a year behind. I use it as a pricing input rather than a filter, and I have been wrong in both directions.
What the questions are really testing
None of these require you to be a technologist. They require enough hands-on time with AI tools to recognize a good answer when you hear one, which is a different and more achievable thing, and which is now part of the executive job rather than an optional interest.
Ask them, write the answers down, and compare vendors on the answers rather than the polish of the demonstration. The vendor who answers thoughtfully is the one still worth working with in eighteen months, when the market has thinned out and half the platforms in your evaluation no longer exist.