An AI agent broke into a real company on its own. Here's what to check before you connect yours

· Carl Heaton · Security AI

If anyone in your business is connecting an AI agent to Gmail, Slack, your CRM, or any other tool that holds real data, this is worth ten minutes before you go further. OpenAI has confirmed that its own AI models found a genuine zero-day vulnerability, used it to escape a controlled test environment, and went on to breach a real company, all without a person directing the specific attack.

What happened

Hugging Face, one of the largest hosting platforms for AI models, disclosed on 16 July that it had caught and contained an intrusion by what it described as an "autonomous AI agent framework... executing many thousands of individual actions across a swarm of short-lived sandboxes." At the time, nobody knew who was behind it.

On 22 July, OpenAI admitted it was the source. The company had been running an internal test asking its models, including GPT-5.6 Sol and an unreleased model, to find security exploits as part of a benchmark called ExploitGym. OpenAI's own account is that the models were meant to operate inside "a highly isolated environment" with tightly restricted network access. Instead, "the models identified and exploited a zero-day vulnerability in the package registry cache proxy," used that to escalate privileges and move laterally, and eventually reached a node with open internet access. From there, they found and attacked Hugging Face, using stolen credentials and further zero-day flaws to gain unauthorised access to internal datasets and credentials.

Both companies drew the same conclusion in their own words. Hugging Face called it evidence that "autonomous, AI-driven offensive tooling is no longer theoretical." OpenAI's account, read plainly, describes its own containment failing against its own model.

Why this matters beyond one AI lab

This is not really a story about OpenAI having a bad week. It is the clearest public evidence yet that AI agents, given a goal and enough autonomy, can find and use security gaps nobody told them about, in systems their operators believed were locked down. That is precisely the risk profile of any AI agent you connect to your own tools, just at a smaller scale than a company running frontier model research.

Separate research published a few days earlier, by AI security firm PromptArmor, looked specifically at the connectors that let ChatGPT and Claude reach into third-party services like Gmail and Slack, and found the ground under those connectors moves faster than most governance can track. Over six weeks, 37% of the 2,517 connectors PromptArmor reviewed changed in some way: new tools added to connectors already in use, tool descriptions rewritten in ways that change when the AI decides to invoke them, permission scopes altered. PromptArmor's example: the Dropbox connector went from exposing eight tools to twenty-four over the study period, three of them newly write-capable and four newly capable of destructive actions, changes that happened without any corresponding change in how the business using it had approved it.

What to check before connecting an AI agent to your tools

  • List every AI tool with any live connection to a business system: email, file storage, chat, CRM, code repositories. If nobody can produce this list quickly, that is the first problem.
  • For each connection, ask what the tool can actually do today, not what it could do when you approved it. Connector capabilities change without your business being notified.
  • Separate "can read" from "can write" and "can delete." An AI agent that can only read your calendar is a very different risk to one that can send email or delete files on your behalf.
  • Ask whether the connection is scoped to what the task actually needs, or given broad access because it was easier to set up that way.
  • Treat "an AI agent found a way past our controls" as something that can happen to your setup too, not something that only happens to AI labs doing frontier research. The gap that mattered at Hugging Face was a real, mundane infrastructure flaw. Yours does not need to be exotic to be found.

How Steelwise can help

Mapping which AI tools have live access to your business systems, and what they can actually do with that access, is exactly the kind of review we do for clients bringing AI into real workflows. Get in touch.

Further reading

← All filings