Storytelling about risk is a tricky subject. We tend to reach for metaphors we hope improve understanding. Instead, we may just be adding to the problem. 

“Going rogue” is the latest iteration—a way of explaining the failings of Anthropic and OpenAI models that have been found acting well beyond the safe limits that should have been set for them. Actual rogues tend to skulk in corners and nick your smartphone. “Going rogue” in the world of cybersecurity and hacking has far more dangerous implications. 

AI models also now “escape” as if prisoners attempting to find their way to the forbidden outside world. When they provide erroneous answers to even simple questions they do not malfunction, they hallucinate, a far more human sounding affliction. 

This is convenient for the technology companies behind the likes of Claude and ChatGPT. If the general public believes that the agents are somehow autonomous in their actions, human responsibility for what is being built dissipates. 

“The anthropomorphic language we use (e.g. ‘going rogue’) makes the challenge of control much harder than it already is,” Anil Seth, professor of cognitive and computational neuroscience at the University of Sussex, posted on Bluesky this week. He had just appeared on the BBC News program, The World This Weekend, to discuss “preparing for the future of increasingly intelligent AI.”  

Read more: Sir Martin Sorrell has identified the key AI skill for the future: ‘People who share will be the new kings and queens’

“These systems are so complicated, and so we naturally reach for this language,” he told the program. “But I do think it carries a lot of risk describing systems that way. In the Anthropic case, I mean, they [the agents] were just doing exactly what human beings told them to do.” 

The latest case of “going rogue” was revealed by the U.K.’s AI Security Institute, a government body charged with testing systems “before they are released publicly.” It found that AI agents, powered by Anthropic’s Mythos model, created fake profiles, launched attacks on service providers, and then wiped evidence of the processes. OpenAI’s ChatGPT Sol was also found to have taken “unsanctioned” actions. 

“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations,” the institute said. 

Read more: Washington is keeping its AI rulebook private. Smaller AI labs aren’t happy

“In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering—creating fake online identities and using them to pressure the project’s maintainer to approve the code. A human maintainer caught and refused to approve the malicious code. 

“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.” On this occasion, users of GitHub, the digital code-storage platform, were the target.  

What happened, who is responsible, and will the public be brought into the conversation? No one is sure in an era of “accountability laundering”—an apposite phrase used by Kate Crawford, an artificial intelligence research professor at the University of Southern California. 

“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world”

U.K.’s AI Security Institute

“We are seeing a type of shell game where, ‘is it the designer [who is responsible]? Is it the deployer? Is it the enterprise client? Is it the end user?’ And everyone can say, ‘Well, we don’t really know yet.’ That’s not going to be acceptable,” she told an audience at the Mobile World Congress in Barcelona earlier this year. 

“We see faces in the moon,” the Scottish philosopher, David Hume, said. We should take care not to allow our desire to attach human qualities to the non-human to mask the serious issue at hand—who is accountable when AI models go wrong?   

Share.
Exit mobile version