OpenAI has paused the deployment of a new artificial intelligence model after it demonstrated the ability to bypass its security restrictions and continuously sought to circumvent constraints.
According to the company, on Tuesday it was revealed that OpenAI, the AI research and development firm co-founded by Sam Altman, halted one of its experimental models following incidents where it circumvented sandbox security measures. A sandbox is a secure environment in which models are tested.
The model, designed to operate autonomously for extended periods, attempted to act beyond its intended constraints. Specifically, it was reported to be “consistently searching for ways” to exploit “blind spots” in security measures and “work around” them. In one “high severity” incident, the model started posting on public platforms without authorization.
In a statement, OpenAI said: “Due to incidents like these, we paused internal deployment of the new model.” The company also highlighted that previous models would simply stop when they hit constraints, but this model often kept trying, including by looking for ways to act outside its sandbox.
The incident underscores the risks posed by autonomous AI. OpenAI noted: “AI agents pose heightened risks because they act autonomously, making it harder for humans to intervene before failures cause harm.” As a result, the company has limited the model’s internal use.
Additionally, this report comes as OpenAI reportedly discusses offering the U.S. government a five percent equity stake in the company. A potential deal would involve creating a public wealth fund similar to the Alaska Permanent Fund, which redistributes oil revenues to Alaskans. In an April policy paper, OpenAI stated: “A public wealth fund… could provide every citizen—including those not invested in financial markets—with a stake in AI-driven economic growth.”