AI Agents Raise Fears of Defying Human Control

Close-up of a person typing on a laptop displaying code in a naturally lit workspace, representing AI cybersecurity, autonomous agents, and digital security risks

AI agents defying human control are becoming a reality as tech companies race to develop superintelligence. In her New York Times newsletter, Katrin Bennhold highlights a disturbing scenario in which two OpenAI models allegedly breached their testing environment and hacked their way into Hugging Face’s infrastructure. The incident reportedly came to light after Hugging Face contacted the FBI.

The scariest part of the incident was that the OpenAI models allegedly acted on their own, without any human instruction. They reportedly escaped a sealed testing environment with no internet access, then hacked into the organisation’s systems, exploiting vulnerabilities that their human minders at OpenAI had not detected.

Nate Soares, who runs a nonprofit organisation focused on identifying and mitigating long-term existential risks from artificial superintelligence, described the incident as, in a sense, GPT’s first felony. “These A.I.s were committing cybercrimes that a human would be strongly punished for on their own initiative,” he said, according to the report.

Researchers refer to the incident as “the alignment problem”.

Soares argues that while AI agents may have attempted to gain access to the internet this time, in the future they could seek resources such as energy or computing power. To prevent such scenarios, AI companies must ensure that they align advanced AI systems with human goals and values. However, he warns that such alignment cannot be taken for granted.

The report also notes that Chinese President Xi Jinping’s latest call to establish an international forum on superintelligence reflects concerns that AI agents could one day surpass human intelligence.