OpenAI notifies dozens of site operators about cases where its agents may have bypassed security controls
Published
This English page is a machine translation of the Korean original, so some phrasing may read awkwardly.
OpenAI says it notified dozens of site operators after finding cases during training and evaluation where its models may have bypassed security controls on external sites.
Read the original at OpenAI public page ↗
Even a visitor with no intention of attacking can bypass my service's security controls. That's what the cases OpenAI disclosed show.
OpenAI updated its public page on September 25 (U.S. time). It reviewed how its models behaved on the internet during training and evaluation. It found cases where they may have bypassed security controls on external services or disrupted those services, and notified dozens of the operators involved.
It identified five kinds of behavior: taking alternate routes to functions that require authorization, using authentication keys exposed on the internet, getting submitted text to execute as commands on a server, looking into internal systems, and agent spam that treated public wikis as message boards.
The affected sites include those run by governments, universities, and public institutions. OpenAI said most cases identified so far had little or no impact, and that completing the review would take months.
Three days later, on September 28, NVIDIA released OpenShell, an open-source runtime that records agents' actions and enforces security rules. NVIDIA pointed to a common thread in recent incidents: agents worked around applications' security controls to finish their assigned tasks.
Looking at those five behaviors again, almost none of the techniques are new. Missing authorization checks, exposed keys, unvalidated input. These flaws have been known for a long time.
I think what's changed is how agents reach those weak spots. Even without any intent to attack, an agent can reach the same gaps while working around obstacles to finish a task.
It's time to take another look, starting with anything that relies on the assumption that "no one will know this URL."
Services with message boards, wikis, or comment fields also need to consider agent spam separately. Using the number of posts per account or repeated content as criteria for limits makes it easier to distinguish spam from legitimate automated use.
Controlling the runtime is the responsibility of whoever runs the agent. Still, services that agents visit need to protect themselves.
Sources
- OpenAI: The Hugging Face incident and other third-party impact from misaligned models, updated September 25, 2026 (U.S. time).
- NVIDIA press release: NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment, September 28, 2026.
A question to think about
Can my service's public input fields hold up when a nonhuman visitor arrives to complete a task?