← All PostsAgentic AI · Board Governance

AI Agents Have Left the Chat

Adrian DunkleySeptember 8, 20269 min read
A dark office lit only by ambient desk lighting, representing the overnight window in which AI agents now run business workloads without a person present

Photo via Unsplash

Definitions, and the numbers

Chat AI answers a question and stops; a person is the loop. Agentic AI is handed an assignment, plans the steps, uses tools such as a browser, a database or a terminal, and returns finished work with nobody watching.

What shipped. Claude Fable 5.1, 1 September 2026. On Terminal-Bench-Science 0.1, an agentic benchmark of 70 tasks run in a terminal, it scored 52.6 percent against 24.7 for Fable 5 and 29.0 for Opus 5. On Terminal-Bench 4.0, 55.8 against 42.0.

The line your CFO should read twice. Headline pricing did not move, at 10 dollars per million input tokens and 50 per million output. Cache reads fell 75 percent to 25 cents. A long-running agent re-reads its context constantly, so cached input dominates an agentic bill in a way it never dominated a chat bill. That is what makes an overnight run affordable.

The catch. 52.6 percent means it failed 47.4 percent of the tasks, on a benchmark whose error margin Anthropic puts at 3.5 to 4.5 percentage points. The trajectory is the story, not the level.

The complication. OpenAI moved the other way, removing agent mode from ChatGPT in early August 2026 with no notice and shutting the Atlas browser on 9 August.

Why the governance maths changes. A junior associate misreads one clause and a partner catches it. An agent applies the same misreading to all 40 contracts in the queue, identically and well formatted. Correlated failure at speed does not resemble a bad hire, which is what most governance frameworks were built around.

Anthropic shipped Claude Fable 5.1 on 1 September 2026. The headline figure in the release is 52.6 percent on Terminal-Bench-Science 0.1, an agentic benchmark of 70 tasks that ask a model to work in a terminal: read a data set, run the computations, submit a result. Fable 5, three months earlier, scored 24.7 percent on the same tasks. Opus 5 scored 29.0. On Terminal-Bench 4.0 the jump was 42.0 to 55.8 percent. Those are not the numbers of a chat product getting more articulate. They are the numbers of a system getting better at finishing work nobody is watching it do.

The line in that release my CFO clients should read twice is not a benchmark at all. Headline pricing did not move, at 10 dollars per million input tokens and 50 per million output. Cache reads dropped 75 percent, from one dollar to 25 cents per million. A long-running agent re-reads the same context over and over as it works through a task, so cached input dominates an agentic bill in a way it never dominated a chat bill. Anthropic did not make the model dramatically cheaper. It made the specific shape of usage that agents produce dramatically cheaper, which is a product decision about what they expect people to run.

Chat answers, agents finish

For most firms, AI over the last two years has meant an employee typing a question into a browser tab and reading what comes back. The employee is the loop. They decide what to ask, they judge the answer, they carry it into the actual work. Everything the model does is bounded by someone's attention span.

An agent is handed an assignment instead of a question. It plans the steps, picks the tools it needs, and works until it has something to hand back. A research report, a reconciliation, a first-pass code change, an operational review. The person sets the brief and reads the deliverable, and the middle is gone. That middle is where most of the labour used to live, which is why the productivity arithmetic changes and why the governance arithmetic changes with it.

What 52.6 percent actually licenses you to conclude

A doubled benchmark reads like a solved problem and this is not one. Fifty-two point six percent means the model failed 47.4 percent of the tasks. Anthropic reports an error margin of roughly 3.5 to 4.5 percentage points on Terminal-Bench-Science 0.1, which is wide enough that the difference between two nearby scores on that benchmark is often not a difference at all. A frontier agent now completes about half of a bounded technical workload without help, up from a quarter three months ago.

Half is enormous compared to a quarter and it is nowhere near enough to remove review. Both are true, and boards tend to hear only whichever one their vendor is selling. The trajectory is the story here, not the level. If the level is what convinces you, you have been shown the wrong chart.

The vendor road is not straight

While Anthropic was tuning a model for long-running autonomy, OpenAI was withdrawing the equivalent capability from its consumer product. Operator, the original browser agent, was folded into ChatGPT agent in July 2025 and shut down on 31 August 2025. Agent mode was then removed from ChatGPT in early August 2026, with no advance deprecation notice and no migration guide. The Atlas browser shut down on 9 August 2026. Users were pointed to ChatGPT Work and a separate cloud browser feature, neither of which does the thing agent mode did: drive a logged-in browser session and hand control back to a human at the login and payment steps.

I am not reading that as OpenAI conceding the category. It is a large company consolidating a product line. But it is a live demonstration of a risk that agentic deployment carries and chat deployment mostly did not. If you build a workflow on an agent that drives a browser, and the vendor removes the agent that drives the browser, you have a broken process and a team who has forgotten how to do it by hand. Chat was replaceable in an afternoon. A workflow is not.

The practical consequence is that the second question in any agent deployment, right after what it does, is what happens to the work if this capability disappears in ninety days. That question was theatre in 2024. It stopped being theatre on 9 August.

What Shipped, and What It Costs

  • 1 September 2026Claude Fable 5.1 and Mythos 5.1 released, generally available across the Claude API, AWS, Google Cloud and Azure
  • 52.6%Terminal-Bench-Science 0.1, against 24.7% for Fable 5 and 29.0% for Opus 5
  • 55.8%Terminal-Bench 4.0, against 42.0% for Fable 5
  • $0.25Per million cached input tokens, down 75% from $1.00; headline input and output pricing unchanged at $10 and $50
  • 1M tokensContext window, with a 128K maximum output
  • 9 August 2026Date OpenAI shut down the Atlas browser, weeks after removing agent mode from ChatGPT without notice

What this looks like in a real firm

The deployments I have been closest to this year are unglamorous and they share a shape. Somebody took a task that was already defined tightly enough to write down, moved it into the overnight window, and put a person on the front of it in the morning instead of throughout the day.

The clearest case is a contract review queue. Agents draft first-pass responses to standard commercial contracts overnight, each redline carrying the reasoning behind it, and the associates arrive to a queue rather than a pile. Turnaround on standard agreements went from days to inside a working day. The firm could do this because it already had a house position on the twenty clauses that come up repeatedly, written down, so the agent had something to be measured against and the associate had something to check the redline with. Firms without that document got a queue of confident, well-formatted, unreviewable drafts.

Reconciliation work has moved the same way, from teams running queries and phoning branches to a nightly run and a morning exception list. So has overnight preparation of analytical drafts, and first-pass triage of flagged transactions before a compliance officer starts their day. The gain in each case shows up in the shape of the morning rather than in headcount. The first hour now starts with a decision instead of with data collection.

I do not have clean before-and-after productivity numbers I would defend in front of a CFO for any of these, and neither does anyone else this early. Boards get hurt pattern-matching on figures like that. What I can say is which ones held up over six months and which got switched off, and the difference was almost never model quality.

Whose name is on it

A junior associate who misreads a clause makes that mistake once, on the contract in front of them, and a partner catches it. An agent that misreads the same clause applies the identical reasoning to all forty contracts in the queue. There is no variation in the error to make it visible, and the output is beautifully formatted whether the thinking was sound or not. Correlated failure at speed is the actual risk profile of agentic deployment, and it does not resemble the risk profile of a bad hire, which is what most governance frameworks were written around.

So the ownership question is narrower than it usually gets asked. Not "who owns AI" but three specific things. Who sets the standard the agent's output is reviewed against. What proportion of output stays sampled once the agent is trusted, and who watches that proportion drift. And who has the authority to pull the agent out of the workflow without a committee. In a mid-sized firm that is the COO or CTO wearing another hat. In a larger one it is an AI Officer with the role written down. Either way it is a person, not a function, and if you cannot name them the deployment is not ready regardless of what the benchmark said.

Boards that have not answered this are carrying an exposure that does not appear anywhere in their current reporting, because nothing has gone wrong yet and the failure mode is quiet until it is forty documents deep.

Three things worth doing this month

Run a frontier agent against the hardest recurring task in your business. Not a demonstration, and not the task your vendor suggests. A real brief, a real deadline, a real deliverable, run in parallel with the person who normally does it, and then read both. That exercise costs a week and it will calibrate your board better than any procurement process, because the evidence comes from your own work rather than from someone else's slide.

Write down who owns the output before the second run, not after the first incident. One name, and the three responsibilities above next to it.

And write down what your firm does if the capability is withdrawn. Which workflows would break, how long the manual fallback takes, and whether anyone still remembers how to run it. OpenAI gave the whole market a free rehearsal of that scenario in August. Use it.

Where this goes

Within a year, running agents on defined recurring processes will be ordinary in firms that compete on cost structure, and the firms that have not started will be roughly a year behind on unit economics rather than behind on technology. The gap will not come from picking the right model. Every serious vendor will have a competent agent, and the models will keep converging the way they have every year since 2023.

It will come from the boring work: knowing which processes are defined tightly enough to hand over, having the written standard that makes an agent's output checkable, and having a named person who is answerable for it. That work is available to any firm in the region right now, at no licence cost, and most of it can be done before you have chosen a vendor at all.

Start with one real task and one named owner. Everything else follows from having done it once.

Related reading across the Caribbean AI network

AI AgentsClaude Fable 5.1Agentic AIBoard GovernanceAI OfficerMaestro AI Labs
About the Author: Adrian Dunkley, The AI Boss

Adrian Dunkley founded StarApple AI, the Caribbean's first AI company, in 2019, and chairs the Caribbean AI Risk Management Council. He also founded Maestro AI Labs and is President of the Caribbean AI Association. He advises Caribbean boards on AI programmes and writes the AI After Dark newsletter for executives and board members.

Connect ↗