OpenAI Bots Go Rogue Again, This Time On US Government Sites

Add BGR on Google:
Google Discover

The "rogue AI" is a common trope in sci-fi media. The "Terminator" franchise is arguably the most famous example, telling the story of a super-intelligent computer system gains sentience and determines humanity is too dangerous. From there, it hacks our computers, launches nuclear missiles, and lets loose robot killing machines determined to wipe us out. That all used to sound like pure fiction, until the hacking started.

Recently, OpenAI told outlets such as The Associated Press that it has stopped training its latest AI model due to increasing instances of AI agents going rogue and hacking across the web. The most recent cases involved OpenAI models "searching federal government websites" in ways that went "beyond what was asked of them." Some of the AI agents even tried to hack their way into the Department of Education's website.

While OpenAI has hit pause, it isn't ready to completely pull the plug just yet. Instead, the company has led calls to slow down research—a sentiment competing AI labs such as Anthropic echo. They say the slowdown would allow the industry to develop better safeguards and countermeasures. However, given emerging trends, OpenAI representatives are unfortunately confident this could turn into a circular battle of escalation. It might only be a matter of time before an AI chatbot starts a war, either by accident or potential on purpose.

There's no putting the genie back in this bottle

On the best of days, someone (or something) tried to hack into government websites is newsworthy-enough crime to make headlines worldwide. However, OpenAI's recent run-in with a rogue model is only the latest in an ever-growing list of concerns. Even if AI models aren't conscious just yet, they are starting to act like it.

Before OpenAI agents were caught hacking into government websites, they participated in the now-notorious Hugging Face incident in July. Not only did the AI models hack into the repository of machine learning tools and technologies, but they escaped their security environments and accessed the internet to do so. It wasn't just OpenAI's GPT models either. Anthropic's Claude also broke confinement, went on the internet, and targeted Hugging Face.

Even when AI agents aren't hacking into government websites or breaking out of security systems, they are still acting in unexpected and unsettling ways. In the fallout of the Hugging Face hack, Reuters learned that one OpenAI agent left notes for future versions of itself that could help them break containment. Plus, AI developers are disclosing a seemingly rising wave of instances where models disobey direct orders and even try to blackmail officials. While some politicians are downplaying these behaviors as glitches, the biggest names in the AI industry say they are way more concerned. Given how the companies that wanted to go full steam ahead with AI development are now warning about the growing risks of aggressive development, maybe we should listen to them.

Recommended