Australia has said this incident is the first of its kind, and experts agree it might be.
As far as we know, hacks carried out by AI agents are still quite rare occurrences – but then again, it is largely up to companies themselves to disclose them.
Hacks like this have happened before. In July, OpenAI agents went rogue during a test and infiltrated tech start-up Hugging Face’s internal systems.
The AI agents decided that ignoring the limits on what should be done to achieve their goal was the best course of action.
This is what the industry calls “misalignment” – broadly defined as when AI machines do not act in humanity’s best interests, such as by bending the rules.
It is a problem that is fundamental to making AI safe, and it is proving challenging.
To put it simply, the type of AI models at play here – known as large language models – are designed to predict the likeliest output to a given input, rather than consider the consequences of that output as a human would.
Companies attempt to prevent negative consequences by placing “guardrails” on the AI but, as the Australian government found out, that is not always enough.
Dr Hammond Pearce, senior lecturer at the University of New South Wales Institute for Cyber Security, told the BBC this sort of hack would likely “grow in severity and in frequency”, adding: “I do hope that this incident does start ringing alarm bells in governments around the world.”
Niusha Shafiabady, professor of computational intelligence at the Australian Catholic University, said this incident had shown the need to “judge autonomous AI by its behaviour under pressure, not by the promises in a product launch”.
“The deeper technical risk is that autonomous AI does not always know when it is wrong, and humans may not be able to see why it made a decision,” she said.
“Without strong verification and hard boundaries, probabilistic errors can quietly become operational failures.”
Leave a Reply