Photo via Unsplash
Own a number, not a project. "We adopted an AI platform" is a purchase order. "We cut sales preparation from four hours to thirty minutes by the end of Q4" is a business outcome. Whoever owns the figure is the AI boss, whatever the title says.
Build something yourself first. Three weekends, not three hours, on a real piece of work. You are learning where the model helps, where it fabricates, where it wastes your time and where it surprises you, and none of that transfers second hand. Vendor reps can tell who has never touched the tools, and they price accordingly.
Move one metric to completion. Five metrics divide the same attention five ways and none of them moves far enough to be visible. Set the baseline before you start, because a baseline reconstructed afterwards is a negotiation rather than a measurement.
Make the vendor show you the tool failing. Every demo shows the product working, which reveals nothing. Watch whether it fails safely: does it flag uncertainty, decline, or produce a confident wrong answer in the same tone as a correct one. Then ask for the last three customer failures.
Write a one-page standard. Four answers: what data goes in a prompt, who approves AI output before a client sees it, where a human must sign off, and what happens when the model is wrong. The page limit is the mechanism. Five pages let an executive convene a group to decide later, and later is usually never.
Where two of these break. Owning a number fails where the real gain is less risk or better quality rather than saved time, and forcing those into a time-saved metric produces a figure nobody believes. The three weekends are often genuinely unavailable, in which case delegate the practice explicitly rather than skip it.
An executive told me last quarter that his firm was "doing AI." I asked what number would be different in December if it worked. He did not have one, and neither did the vendor whose contract he was about to sign. I have had a version of that exchange in most boardrooms I have sat in this year.
The five practices below are what the executive who owns the programme does differently from the one who does not. None of them is about choosing a vendor.
Own a number
The executives running programmes that work do not say they are doing AI. They say they are cutting sales preparation from four hours to thirty minutes by the end of Q4. AI is the means to that figure, and whoever owns the figure is the AI boss regardless of what the title says.
This changes four things. It changes how you evaluate a vendor, because you now have a test they either pass or fail. It changes how you brief the team, because there is a target rather than a technology. It changes budget allocation, because spend can be traded against expected movement. And it changes what goes to the board: "we moved cycle time on X from Y to Z" instead of "we adopted an AI platform."
Where this breaks is on the outcomes that matter most and measure worst. Fewer contract errors. Better-calibrated credit decisions. Fraud caught two days earlier. Every one of those is worth more than a cycle-time saving and none of them produces a clean before-and-after figure inside one quarter. Executives who force those into a time-saved metric to satisfy this rule end up with a number nobody believes, including them. The alternative is slower: name the outcome, name how you will sample it, and accept that the evidence will be less quotable. Do that rather than manufacture a figure.
Use the tools yourself, seriously, before you approve anything
An executive who has never spent a weekend building something with Claude, ChatGPT, NotebookLM or Cursor cannot judge a vendor demo honestly. This is not a point about credibility. The vendor reps know which executives have never touched the tools, and they pitch and price accordingly, because a buyer who cannot tell a hard problem from an easy one will pay the same for both.
The threshold is measured in weekends, not hours. Three of them, building something that matters: a customer segmentation, a competitive analysis, a first draft of your Q3 strategy note, a rewrite of your standard client proposal. Pay attention to four things while you do it. Where the model genuinely helps. Where it fabricates. Where it wastes your time. And where it surprises you, which is the one people underweight and the one that tells you what is actually new.
The obvious objection is that a chief executive does not have three weekends, and that is often true rather than an excuse. Then delegate the practice rather than skip it: name someone who does have the time, give them the same brief, and put their read on the vendors into the decision record. What is not available is approving a budget on the strength of a demonstration that nobody in the firm is equipped to interrogate, which is where most boards currently are.
One metric this month, not five
The most common way I see AI programmes fail has nothing to do with the vendor or the model. It is five metrics running in parallel. Each gets a fifth of the attention. None moves far enough to be visible. At the end of the quarter the programme has nothing to show and the next budget conversation is harder than the last one.
Pick one. Cycle time on a specific task, or hours saved by a specific person. Set the baseline before you start, because a baseline reconstructed afterwards is a negotiation rather than a measurement. Ship AI at that number until it moves, then pick the next one. Finishing the first is what makes the second easier and the third nearly routine, and it is the credibility that gets both funded.
Ban the demo that only shows the tool winning
Every vendor demonstration shows the product doing something impressive. That reveals nothing, because a product that could not be made to look good in a controlled demonstration would not have reached the market. What you learn nothing about is the shape of the failure.
Ask to see it fail. Not a hypothetical, a live attempt on an input designed to be hard. What you are watching for is whether it fails safely: does it flag uncertainty, does it decline, or does it produce a confident wrong answer in the same tone as a correct one. Then ask what the last three customer failures were and what was done about them. A vendor who says there have been none is either not talking to their customers or not talking straight to you, and both answers should end the meeting.
The one-page standards document
This is the practice with the widest gap between how boring it sounds and how much it decides, so it gets the most space here.
Not a policy, not a governance framework, not a working group's output. One page, answering four questions. What data goes into a prompt and what does not. Who approves AI-generated output before it reaches a client. At what points a human must sign off before anything ships. And what the standard response is when the model turns out to be wrong, including who gets told and how quickly.
The single-page limit is the mechanism rather than a preference about length. A page forces the executive to decide. Five pages allow the executive to convene a group that will decide later, and later is usually never, because by the time the group reports the deployment has been running for two quarters and the standard is now a description of what people already do rather than a constraint on it. One-page documents get read, argued over in the meeting they are handed out in, acted on, and revised. Five-page ones get filed.
Two things I have learned from watching these get written. The data question is the one people get wrong, because they answer it in categories ("no confidential data") that nobody can apply at 4pm with a deadline, rather than in examples that a person can pattern-match against in ten seconds. And the fourth question, what happens when the model is wrong, is almost always left blank, which means the firm's real policy is that whoever notices first decides what to do, and the firm finds that out during an incident.
If all four answers do not fit on a page, you do not have a standard yet. You have a list of things you intend to think about.
The Five, in Short
- A numberOwn a measurable outcome, not a project. Name the figure and the date
- Three weekendsMinimum hands-on build time before approving a budget, or an honest delegation of it
- One metricMoved to completion, with a baseline set before the work starts
- The failure demoWatch the tool break, and ask for the last three customer failures
- One pageData in prompts, output approval, sign-off points, and what happens when it is wrong
What ties them together
None of these five requires technology the firm does not already have access to. Four of them cost nothing but executive attention, and the fifth costs an afternoon. What they require is the specific discipline of owning an outcome rather than owning a project, which is harder than it sounds precisely because owning a project is more comfortable and looks identical in a board pack.
Owning an outcome is the actual work of the role. The vendors, the tools, the frameworks and the strategy documents are all downstream of it, and none of them will supply it.
Start with the number. If you cannot say what will be different in December, nothing else on this list will save the programme.