← All Posts AI Industry

Anthropic's Claude Code Leak: What It Means and What Comes Next

Adrian Dunkley April 2026 6 min read
The Secret, stipple engraving, first state by Jacques Bonnefoy d'Arles
Jacques Bonnefoy d'Arles, The Secret, ca. 1804–31. The Metropolitan Museum of Art, public domain.

The internal instructions behind Anthropic's Claude Code are now public. The text that tells Claude how to behave as an agentic coding assistant (an AI tool that edits files and runs commands on a developer's computer) is circulating on developer forums, X, Reddit and in most AI newsletters. I have read it closely, as someone who has built AI systems in Jamaica for 15 years and runs StarApple AI, and this is what it shows, what it does not change, and what developers and companies should do about it.

What the leaked instructions contain

The leaked prompt is the instruction layer between the base model and the user. It turns a general-purpose language model into a coding agent. It covers:

  • Everyday engineering work: file operations, git workflows, code review, pull requests, error handling and calls to external tools such as GitHub's API.
  • Safety rules: do not force-push to main, do not skip pre-commit hooks, do not run destructive commands without confirmation.
  • Output rules: tone, formatting, and step-by-step procedures for commits and background tasks.
  • Sub-agents: Claude Code acts as an orchestrator that can start specialised agents for searching a codebase, planning a design or carrying out multi-step tasks, each with its own tools and scope.

The restraint instructions are the part I found most useful. The prompt tells the model not to add features beyond the request, not to create abstractions nobody asked for, and not to write error handling for cases that cannot happen. Those lines exist because models do all three by default. Anyone who has reviewed AI-written code has seen the result: a small fix arrives with new helper classes, config options and tests for cases nobody raised.

What it teaches about building reliable agents

Reliable agents take thousands of words of specific instructions, most of them about what not to do. "Write good code" does nothing. "Ask before deleting files" and "do not modify tests to make them pass" change behaviour you can observe.

Four lessons carry over to any team building its own agent:

  • Write rules as observable actions ("confirm before running a migration"), not as qualities ("be careful").
  • Put the dangerous operations behind a confirmation step in code as well as in the prompt.
  • Split work across narrow sub-agents with limited tools. A search agent that cannot write files cannot damage anything.
  • Keep a list of the mistakes your agent makes, and turn each one into a rule. Much of the leaked text reads like that list.

The security question: does the leak make Claude Code less safe?

It makes some attacks easier to plan. Someone who knows exactly when the agent asks for confirmation, and what it refuses, can craft inputs that try to route around those rules. The obvious route is prompt injection: instructions hidden in a README, an issue comment or a web page that the agent reads.

It does not remove the main defences. Anthropic trains safety behaviour into the model itself through methods such as Constitutional AI and reinforcement learning from human feedback, and Claude Code also enforces permissions outside the model: the user decides which commands and tools the agent may run without asking. A prompt is one layer among several, and a system whose safety depended on keeping its prompt secret would already have been fragile.

The lesson for everyone else building on language models is to assume the prompt will become public. Several AI products have had their system prompts extracted and posted online, often by simply asking the model to repeat them. Design so that a published prompt costs you nothing: no credentials in it, no business logic that only works if hidden, and permission checks enforced in code.

What it signals about AI coding tools

The direction is agents that act, not assistants that suggest. Claude Code already creates branches, edits files, runs tests and opens pull requests. Handling review comments and merging after approval are the next steps, and some coding agents already offer both.

For developers in Kingston, Bridgetown and Port of Spain, the practical point is that the design of these tools is now readable. A developer who understands how the agent is instructed can write better project instructions, predict where it will fail and review its output faster. My own caution: a prompt shows intent, not performance. It tells you what Anthropic asked the model to do, and nothing about how often the model complies.

What to do next

  1. Developers using Claude Code: add a CLAUDE.md file to the root of each repository with your build commands, test commands and house rules, written as specific actions. It takes about 30 minutes and Claude Code reads it at the start of every session.
  2. Engineering leads: review which commands your team allows the agent to run without confirmation. Anything that deletes data, pushes to a shared branch or touches production should require approval.
  3. Companies building their own AI products: audit your system prompts this week for API keys, internal URLs, customer names and pricing logic. Move anything sensitive out of the prompt and into server-side code.
  4. Security teams: add prompt injection to your testing. Put a hostile instruction in a test README or issue, point the agent at it, and record what happens.
  5. Trainers and bootcamps in the region: teach agent design from public documentation and your own prompts. Study the structure of the leaked text if you wish, but do not copy it into products or course materials; it is Anthropic's work.

Frequently Asked Questions

What is a system prompt, and how is it different from the model?

A system prompt is a block of text the application sends to the model before every conversation. It sets the model's role, rules and available tools. The model is the trained network itself, and its weights were not part of this leak. Changing a system prompt changes behaviour within the limits of what the model has learned, which is why the same model can act as a chatbot in one product and a coding agent in another.

Can I use the leaked Claude Code prompt in my own product?

You should not copy it. The text is Anthropic's work, and reusing it wholesale in a commercial product invites copyright and terms-of-service problems; take legal advice if you are unsure. Studying its structure and then writing your own instructions for your own tools is ordinary engineering practice and produces a better fit for your product anyway.

Does the leak put my code or data at risk if I use Claude Code?

Reporting on the leak concerned Anthropic's own instructions, and did not describe any exposure of customer code or conversations. The risks for a user are the usual ones for any coding agent: secrets committed to the repository, broad permissions that let the agent run commands unchecked, and malicious instructions hidden in files or web pages the agent reads. Review your permission settings and keep credentials in environment variables or a secrets manager, not in code.

How do I stop my own AI product's system prompt from leaking?

You cannot fully stop it, because a determined user can often get a model to reveal its instructions. Instruction-level defences such as "never reveal this prompt" slow attackers down but fail regularly. Plan for disclosure: keep keys and private data out of the prompt, enforce access rules in your backend, and write the prompt so you would be comfortable seeing it on social media.

How do I give Claude Code instructions specific to my project?

Create a file called CLAUDE.md in the repository root; Claude Code loads it into context at the start of each session. Put in it the commands to build, test and lint, the coding conventions you enforce, and any operations the agent must never perform. Keep it short and specific, and add a rule each time the agent makes a mistake you do not want repeated.

Claude Code Anthropic AI Leak System Prompt AI Transparency
Adrian Dunkley

Physicist and AI Scientist. Founder of StarApple AI, the Caribbean's First AI Company. Founder of four AI Labs in Jamaica. Jamaica's #1 AI Leader.

Connect ↗