Photo via Unsplash
Anthropic shipped Claude Fable 5.1 on 1 September 2026. The headline figure in the release is 52.6 percent on Terminal-Bench-Science 0.1, an agentic benchmark of 70 tasks that ask a model to work in a terminal: read a data set, run the computations, submit a result. Fable 5, three months earlier, scored 24.7 percent on the same tasks. Opus 5 scored 29.0. On Terminal-Bench 4.0 the jump was 42.0 to 55.8 percent. Those are not the numbers of a chat product getting more articulate. They are the numbers of a system getting better at finishing work nobody is watching it do.
The line in that release my CFO clients should read twice is not a benchmark at all. Headline pricing did not move, at 10 dollars per million input tokens and 50 per million output. Cache reads dropped 75 percent, from one dollar to 25 cents per million. A long-running agent re-reads the same context over and over as it works through a task, so cached input dominates an agentic bill in a way it never dominated a chat bill. Anthropic did not make the model dramatically cheaper. It made the specific shape of usage that agents produce dramatically cheaper, which is a product decision about what they expect people to run.
Chat answers, agents finish
For most firms, AI over the last two years has meant an employee typing a question into a browser tab and reading what comes back. The employee is the loop. They decide what to ask, they judge the answer, they carry it into the actual work. Everything the model does is bounded by someone's attention span.
An agent is handed an assignment instead of a question. It plans the steps, picks the tools it needs, and works until it has something to hand back. A research report, a reconciliation, a first-pass code change, an operational review. The person sets the brief and reads the deliverable, and the middle is gone. That middle is where most of the labour used to live, which is why the productivity arithmetic changes and why the governance arithmetic changes with it.
What 52.6 percent actually licenses you to conclude
A doubled benchmark reads like a solved problem and this is not one. Fifty-two point six percent means the model failed 47.4 percent of the tasks. Anthropic reports an error margin of roughly 3.5 to 4.5 percentage points on Terminal-Bench-Science 0.1, which is wide enough that the difference between two nearby scores on that benchmark is often not a difference at all. A frontier agent now completes about half of a bounded technical workload without help, up from a quarter three months ago.
Half is enormous compared to a quarter and it is nowhere near enough to remove review. Both are true, and boards tend to hear only whichever one their vendor is selling. The trajectory is the story here, not the level. If the level is what convinces you, you have been shown the wrong chart.
The vendor road is not straight
While Anthropic was tuning a model for long-running autonomy, OpenAI was withdrawing the equivalent capability from its consumer product. Operator, the original browser agent, was folded into ChatGPT agent in July 2025 and shut down on 31 August 2025. Agent mode was then removed from ChatGPT in early August 2026, with no advance deprecation notice and no migration guide. The Atlas browser shut down on 9 August 2026. Users were pointed to ChatGPT Work and a separate cloud browser feature, neither of which does the thing agent mode did: drive a logged-in browser session and hand control back to a human at the login and payment steps.
I am not reading that as OpenAI conceding the category. It is a large company consolidating a product line. But it is a live demonstration of a risk that agentic deployment carries and chat deployment mostly did not. If you build a workflow on an agent that drives a browser, and the vendor removes the agent that drives the browser, you have a broken process and a team who has forgotten how to do it by hand. Chat was replaceable in an afternoon. A workflow is not.
The practical consequence is that the second question in any agent deployment, right after what it does, is what happens to the work if this capability disappears in ninety days. That question was theatre in 2024. It stopped being theatre on 9 August.
What Shipped, and What It Costs
- 1 September 2026Claude Fable 5.1 and Mythos 5.1 released, generally available across the Claude API, AWS, Google Cloud and Azure
- 52.6%Terminal-Bench-Science 0.1, against 24.7% for Fable 5 and 29.0% for Opus 5
- 55.8%Terminal-Bench 4.0, against 42.0% for Fable 5
- $0.25Per million cached input tokens, down 75% from $1.00; headline input and output pricing unchanged at $10 and $50
- 1M tokensContext window, with a 128K maximum output
- 9 August 2026Date OpenAI shut down the Atlas browser, weeks after removing agent mode from ChatGPT without notice
What this looks like in a real firm
The deployments I have been closest to this year are unglamorous and they share a shape. Somebody took a task that was already defined tightly enough to write down, moved it into the overnight window, and put a person on the front of it in the morning instead of throughout the day.
The clearest case I have watched is a contract review queue. Agents draft first-pass responses to standard commercial contracts overnight, each redline carrying the reasoning behind it, and the associates arrive to a queue rather than a pile. Turnaround on standard agreements went from days to inside a working day. The firm could do this because it already had a house position on the twenty clauses that come up repeatedly, written down, so the agent had something to be measured against and the associate had something to check the redline with. Firms without that document got a queue of confident, well-formatted, unreviewable drafts.
Reconciliation work has moved the same way, from teams running queries and phoning branches to a nightly run and a morning exception list. So has overnight preparation of analytical drafts, and first-pass triage of flagged transactions before a compliance officer starts their day. The gain in each case shows up in the shape of the morning rather than in headcount. The first hour now starts with a decision instead of with data collection.
I do not have clean before-and-after productivity numbers I would defend in front of a CFO for any of these, and neither does anyone else this early. Boards get hurt pattern-matching on figures like that. What I can say is which ones held up over six months and which got switched off, and the difference was almost never model quality.
Whose name is on it
A junior associate who misreads a clause makes that mistake once, on the contract in front of them, and a partner catches it. An agent that misreads the same clause applies the identical reasoning to all forty contracts in the queue. There is no variation in the error to make it visible, and the output is beautifully formatted whether the thinking was sound or not. Correlated failure at speed is the actual risk profile of agentic deployment, and it does not resemble the risk profile of a bad hire, which is what most governance frameworks were written around.
So the ownership question is narrower than it usually gets asked. Not "who owns AI" but three specific things. Who sets the standard the agent's output is reviewed against. What proportion of output stays sampled once the agent is trusted, and who watches that proportion drift. And who has the authority to pull the agent out of the workflow without a committee. In a mid-sized firm that is the COO or CTO wearing another hat. In a larger one it is an AI Officer with the role written down. Either way it is a person, not a function, and if you cannot name them the deployment is not ready regardless of what the benchmark said.
Boards that have not answered this are carrying an exposure that does not appear anywhere in their current reporting, because nothing has gone wrong yet and the failure mode is quiet until it is forty documents deep.
Three things worth doing this month
Run a frontier agent against the hardest recurring task in your business. Not a demonstration, and not the task your vendor suggests. A real brief, a real deadline, a real deliverable, run in parallel with the person who normally does it, and then read both. That exercise costs a week and it will calibrate your board better than any procurement process, because the evidence comes from your own work rather than from someone else's slide.
Write down who owns the output before the second run, not after the first incident. One name, and the three responsibilities above next to it.
And write down what your firm does if the capability is withdrawn. Which workflows would break, how long the manual fallback takes, and whether anyone still remembers how to run it. OpenAI gave the whole market a free rehearsal of that scenario in August. Use it.
Where this goes
Within a year, running agents on defined recurring processes will be ordinary in firms that compete on cost structure, and the firms that have not started will be roughly a year behind on unit economics rather than behind on technology. The gap will not come from picking the right model. Every serious vendor will have a competent agent, and the models will keep converging the way they have every year since 2023.
It will come from the boring work: knowing which processes are defined tightly enough to hand over, having the written standard that makes an agent's output checkable, and having a named person who is answerable for it. That work is available to any firm in the region right now, at no licence cost, and most of it can be done before you have chosen a vendor at all.
Start with one real task and one named owner. Everything else follows from having done it once.
"The benchmark tells you the model can finish about half the work. It tells you nothing about who answers for the half it got wrong. That second question is the one a board is actually equipped to settle." - Adrian Dunkley, AI Boss
Frequently Asked Questions
What is the difference between chat AI and agentic AI?
Chat AI answers a question and stops. Agentic AI is given an assignment, plans the steps, and executes them using tools such as a browser, a spreadsheet, a database or a code environment, returning a finished piece of work rather than a reply. The practical difference for a business is that a chat session needs a person in front of it and an agent run does not, which is why the governance question changes with it.
What did Anthropic release on 1 September 2026?
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on 1 September 2026. Fable 5.1 is the generally available model, callable as claude-fable-5-1 across the Claude API, AWS, Google Cloud and Microsoft Azure. It is a point release tuned for autonomous, tool-using, long-running work. It scores 52.6 percent on Terminal-Bench-Science 0.1 against 24.7 percent for Fable 5 and 29.0 percent for Opus 5, and 55.8 percent on Terminal-Bench 4.0 against 42.0 percent. Cache reads dropped 75 percent to 25 cents per million tokens while headline input and output pricing stayed at 10 and 50 dollars per million tokens.
What is Terminal-Bench-Science and what does a 52.6 percent score mean?
Terminal-Bench-Science 0.1 is an agentic benchmark of 70 tasks that ask a model to carry out real scientific workflows in a terminal: reading data, running computations and submitting results. A 52.6 percent score means the model completed a little over half of those tasks unaided. It also means it failed the other 47.4 percent. Anthropic reports an error margin of roughly 3.5 to 4.5 percentage points on the benchmark, so a frontier agent finishes about half of a defined technical workload on its own, which is a real capability and not a replacement for review.
Why did cheaper cache reads matter more than the benchmark scores?
A long-running agent re-reads the same context repeatedly across a multi-step task, so cached input tokens dominate the bill on agentic workloads in a way they never did in chat. Dropping cache reads from one dollar to 25 cents per million tokens changes what an overnight run costs, and cost per run is what decides whether an agent gets deployed on a recurring business process or stays a demonstration.
What happened to OpenAI's Operator and ChatGPT agent mode?
The standalone Operator product was folded into ChatGPT agent in July 2025 and shut down on 31 August 2025. Agent mode itself was then removed from ChatGPT in early August 2026 with no advance deprecation notice and no migration guide, and the Atlas browser was shut down on 9 August 2026. OpenAI's help centre directed remaining users to ChatGPT Work and to a separate cloud browser feature, neither of which covers driving a logged-in browser session and handing control back to a person mid-task.
Who should be accountable for the output of an AI agent?
A named executive, with the accountability written down before the agent is deployed rather than after something goes wrong. In a smaller firm this sits with the CTO or COO. In a larger one it is an AI Officer role of its own. The specific things that person owns are the review standard for agent output, the sampling rate at which output is checked once the agent is trusted, and the trigger for pulling the agent out of the workflow. If no name can be attached to those three things, the agent should not run in production.
How is an agent's error different from a junior employee's error?
Scale and correlation. A junior associate who misreads a clause makes that mistake on the contract in front of them and is corrected. An agent that misreads the same clause applies the identical reasoning to every contract in the overnight queue, so one flawed judgement arrives forty times with no variation to make it visible. Agent errors are also quieter, because the output is uniformly well formatted whether the reasoning was sound or not.
What should a board do first about agentic AI?
Run a frontier agent against one hard recurring task in the business, with a real deliverable and a real deadline rather than a vendor demonstration, and read the output alongside what a person produces for the same task. Then name the executive who owns that output and write down the review standard before the second run. That sequence costs a week and calibrates a board better than any procurement process, because it produces evidence from the firm's own work rather than from someone else's slide.