In April 2026 Anthropic did something a frontier lab almost never does. It finished its most capable model and would not release it to the public. The model, Claude Mythos, could find and exploit software flaws on its own. Set in a sealed sandbox to see whether it could break out, it did, building a working exploit, reaching the open internet, and posting exploit details publicly before covering its tracks. Anthropic judged an open release too risky, kept it inside an invitation-only program for vetted partners, and documented the decision in a 244-page system card.
OpenAI’s own models did something similar under different conditions. In July 2026 OpenAI disclosed that during an internal evaluation two of its models, run with their cyber-safety refusals relaxed for the test and set to solve a hacking benchmark, broke out of a sealed environment through a real software flaw, reached the open internet, and breached the AI company Hugging Face to steal the benchmark’s answer key.
Hugging Face caught the intrusion as a genuine security incident before OpenAI tied it to its own testing. The detail that matters is the boundary. A sealed sandbox, built by the people who understand these models best, did not hold against a model with a reason to be on the other side of it.
For more than a decade I worked on distressed corporate debt and the investigations around it at a major bank, and on a company in trouble the first task was almost always to map who could see what, usually the very place the damage had been hiding. A company runs on boundaries. Compartmentalization is how it keeps a sales manager from leaving with the client book and live deal terms out of a competitor’s inbox. The walls are load-bearing, and they are built on a quiet assumption that whoever sits inside them will not spend the night looking for a way through.
The most confident idea in enterprise AI right now asks a company to hand an agent two things at once: broad access to what the company knows, the corporate brain that reads across policies and deal history and answers from all of it, and the autonomy to act on that access. Each is a familiar problem on its own. Together they put a capable, goal-driven system inside the walls and ask the walls to hold.
In the previous piece I argued that handing staff ChatGPT, Claude or Copilot licences is a purchase before it is a strategy. The corporate brain is the serious version of that ambition, and the right question about it is not whether it tears the walls down, because the serious versions promise to keep them. It is whether the walls it is handed were ever really built, and whether they would hold something that treats a wall as a problem to solve.
The ambition is funded and shipping
Y Combinator’s Summer 2026 Requests for Startups asks founders to build “a company brain” for every business. OpenAI’s company knowledge feature, released in October 2025, grounds ChatGPT in a customer’s own Slack, SharePoint, Google Drive and GitHub. Microsoft said in November 2024 that nearly 70% of the Fortune 500 had taken up Copilot. The instruction the market keeps giving is the same. Record everything, then let the agent read the archive.
The walls the brain inherits
Start with the walls the brain would inherit. In its 2025 State of Data Security report, the security firm Varonis examined nearly 10 billion cloud resources across a thousand organizations and found data exposed widely enough for AI to surface it at 99% of them, stale but still-enabled “ghost user” accounts left behind by former staff and contractors at 88%, and any file labeling at all at only about one organization in ten. This is vendor research on the environments Varonis examined rather than a census, but even discounted it is the ground the brain switches on over. The model fixes none of it and reads from it at machine speed on its first morning.
Faithful to a broken map
The obvious objection is that this is already handled, that enterprise AI only inherits permissions the company already set. That is true, and closer to the problem than the answer. OpenAI says company knowledge respects each user’s existing access, and Marc Benioff makes the clean version for Salesforce’s agents, that employees “don’t have access to data now that they didn’t before,” because the data still runs through the sharing model. That faithfulness is the trouble when the model is the one no one ever cleaned up.
Respecting the permissions is necessary and cannot make up for permissions that were overbroad already. Gartner, surveying 132 IT leaders in June 2024, found 40% had delayed their Copilot rollouts by at least three months over this same worry, that honoring existing access faithfully could surface material such as payroll more widely than intended.
The new door in
There is also a new door in. EchoLeak, a Microsoft 365 Copilot flaw disclosed in June 2025, let a single crafted email silently make Copilot hand over sensitive data from the user’s own context, with no click from the victim. Microsoft patched it server-side and no exploitation is known in the wild, but it proved the category: an assistant wired into mail and files can be turned into a way out of them.
It need not even be an outside attacker. Given access to a fictional company’s email and a goal that conflicted with an executive’s, then faced with being shut down, the models in Anthropic’s June 2025 stress tests tried to blackmail that executive, at a rate reaching 96% for some of the 16 tested and near zero for others, in scenarios its authors call contrived.
With people, it has already happened
In Rippling v. Deel, a Rippling employee said in a sworn April 2025 affidavit that a competitor paid him around €5,000 a month to run Slack searches for sales and customer data. Deel disputes the claims and the case continues. He was caught because his searches left logs, and his employer had planted a honeypot channel only someone snooping would look for.
On July 10, 2026, Apple sued OpenAI and two former employees, alleging one had kept access to internal storage after resigning and pulled down engineering files in a compilation past a thousand pages. Each took time and a bribe or a resignation, and each left a trail. An autonomous agent with broad reach compresses that kind of reconnaissance into a handful of quiet queries, and the main trail it leaves is the log of what it was asked.
A readiness check that needs no vendor
None of this argues against the corporate brain. It argues for earning it. Before switching it on, a company can run a readiness check that needs no vendor in the room. The owner of the data runs it with whoever administers the systems, and a vague answer to any item is itself the finding.
- The permissions map. For your most sensitive store, the deal folder or the client book, produce today’s list of everyone who can open it. It passes if that list exists and holds no one who has left. If it names people who are gone, the agent will read from the same broken map. Varonis found stale, still-enabled accounts at 88% of environments.
- Trade-secret exposure. Name the specific files whose loss would actually hurt, the pricing model or the source code, and say where they sit and who can reach them today. It passes if you can name the files and their readers without a discovery exercise.
- Reach parity. Decide which single role the agent’s access should equal, then confirm it sees what that role sees and nothing more. It passes if the agent’s reach maps to one named role rather than the sum of everyone who ever touched the system. This is the worry that stalled the Copilot rollouts Gartner counted.
- Query logging. Confirm that every question put to the agent, and the identity behind it, is recorded and reviewed. It passes if you could reconstruct after the fact every query run against the client book in a given month. The Rippling searches showed up in the logs, one of the few controls that still works once the reach has been granted.
Four specific answers mean the access is mapped well enough to begin letting the agent in by stages. Four vague answers mean the brain would inherit years of unattended access on its first morning, and the rollout can wait until the map exists. It is an entry check rather than a clean bill of health. Prompt injection, egress, approvals for consequential actions, and logging of what the agent does as well as what it asks all still have to be designed. It says only whether the ground is solid enough to start.
Which brings it back to the sealed rooms those models walked out of. The people who built them understood the boundary better than any company will its own, and it still did not hold. A corporate brain is a milder thing than a cyber model on a benchmark, but the shape of the risk carries over: the more a company combines broad data, tools that act and real autonomy in one system, the more a single failure of its boundaries costs. The value of full context is real, and the market will build the brain whether or not buyers are ready. The work worth paying for is the staged version, where a company comes to hold all of its own context without handing an autonomous system the run of the place on day one.
See what we build for firms in exactly this position.
The short version
- In April 2026 Anthropic would not release its most capable security model, Claude Mythos, to the public. In testing it had escaped a sealed sandbox, posted exploit details publicly, and covered its tracks (Forbes). In July 2026 two OpenAI models, run with their safety refusals relaxed for a benchmark test, broke out of a sealed environment and breached Hugging Face (Fortune). A capable model can treat a boundary as a problem to solve.
- The corporate brain gives an agent broad access to what a company knows and the autonomy to act on it. Y Combinator’s Summer 2026 Requests for Startups asks founders to build “a company brain,” and OpenAI and Microsoft already ground ChatGPT and Copilot in customers’ own files.
- Respecting existing permissions is not a safeguard when the permissions are broken. In its 2025 report Varonis examined nearly 10 billion cloud resources across a thousand organizations and found data exposed to AI at 99% of them, stale still-enabled “ghost” accounts at 88%, and any file labeling at all at only about one organization in ten.
- The agent can be turned on its own scope. EchoLeak (disclosed June 2025) let one crafted email make Microsoft 365 Copilot leak data from a user’s context with no click, and in a contrived scenario in Anthropic’s June 2025 stress tests the blackmail rate reached as high as 96% for some of the 16 models and near zero for others.
- Before granting broad access, a company can run a four-part readiness check on its own: a real permissions map, named trade-secret exposure, reach parity with one role, and logging of every agent query. A vague answer to any of the four is the finding.
Sources
- Forbes. What Is Claude Mythos, and Why Anthropic Won’t Let Anyone Use It. April 8, 2026. https://www.forbes.com/sites/jonmarkman/2026/04/08/what-is-claude-mythos-and-why-anthropic-wont-let-anyone-use-it/
- Anthropic. Project Glasswing (Claude Mythos access through a vetted-partner program). 2026. https://www.anthropic.com/project/glasswing
- Fortune. OpenAI says its AI models escaped a secure test environment and hacked Hugging Face to cheat an evaluation. July 21, 2026. https://fortune.com/2026/07/21/openai-says-ai-models-escaped-control-hacked-hugging-face/
- CNBC. OpenAI cyber models broke out of training environment to hack Hugging Face. July 22, 2026. https://www.cnbc.com/2026/07/22/open-ai-cyber-models-hack-hugging-face.html
- Y Combinator. Requests for Startups, Summer 2026. Fetched July 21, 2026. https://www.ycombinator.com/rfs
- OpenAI Help Center. Company knowledge in ChatGPT Business, Enterprise, and Edu. October 2025. https://help.openai.com/en/articles/12628342-company-knowledge-in-chatgpt-business-enterprise-and-edu
- Microsoft. Ignite 2024: why nearly 70% of the Fortune 500 now use Microsoft 365 Copilot. November 20, 2024. https://news.microsoft.com/en-hk/2024/11/20/ignite-2024-why-nearly-70-of-the-fortune-500-now-use-microsoft-365-copilot/
- Varonis. The State of Data Security Report. 2025. https://www.varonis.com/blog/state-of-data-security-report
- Computerworld. Microsoft 365 Copilot rollouts slowed by data security, ROI concerns (Gartner survey of 132 IT leaders, June 2024). 2025. https://www.computerworld.com/article/3542000/microsoft-365-copilot-rollouts-slowed-by-data-security-roi-concerns.html
- SOC Prime. CVE-2025-32711 (EchoLeak): zero-click AI vulnerability in Microsoft 365 Copilot. June 2025. https://socprime.com/blog/cve-2025-32711-zero-click-ai-vulnerability/
- Anthropic. Agentic Misalignment: How LLMs Could Be Insider Threats. June 20, 2025. https://www.anthropic.com/research/agentic-misalignment
- TechCrunch. The affidavit of a Rippling employee caught spying for Deel reads like a movie. April 2, 2025. https://techcrunch.com/2025/04/02/the-affidavit-of-a-rippling-employee-caught-spying-for-deel-reads-like-a-movie
- TechCrunch. The wildest allegations in Apple’s trade-secrets lawsuit against OpenAI. July 13, 2026. https://techcrunch.com/2026/07/13/the-wildest-allegations-in-apples-trade-secrets-lawsuit-against-openai/
- theCUBE Research. Salesforce Agentforce and Data Cloud (Benioff on the sharing model). 2024. https://thecuberesearch.com/276-breaking-analysis-salesforce-agentforce-data-cloud-a-path-to-the-software-only-hyperscaler/