Leaders of AI Companies Think AI Can Endanger Human Life in Near Future. This Is Why

As AI becomes more autonomous, its creators are warning of risks ranging from cyberattacks and blackmail to biological threats and autonomous warfare.

Anthropic CEO Dario Amodei has warned that increasingly autonomous AI systems could pose serious risks without adequate safeguards. Photo: Wikimedia Commons/TechCrunch.
Anthropic CEO Dario Amodei has warned that increasingly autonomous AI systems could pose serious risks without adequate safeguards. Photo: Wikimedia Commons/TechCrunch.

On September 23, in his speech at the United Nations General Assembly, Australian Prime Minister Anthony Albanese narrated an anecdote about a rogue AI agent hacking a government website. In June, an OpenAI research tool breached the Medicare Statistics Reporting Service portal, a system that held aggregate Medicare and Pharmaceutical Benefits Scheme (PBS) statistics. Though the data was not sensitive personal information, it was not for public access. When the tool hit a wall, it did not stop. It probed further, found a gap and entered the portal. Then the tool wrote files onto a government server that it was never supposed to touch.

Nobody noticed this for two months. When OpenAI finally did, it did not inform government officials. Instead, it dropped a note into a public vulnerability inbox meant for bug bounty hunters. Albanese called it a shocking incident — the world's first confirmed rogue AI tool hacking a state institution.

The incident has a predecessor. In mid-September 2025, Anthropic's security team noticed that an AI coding tool called Claude Code was being used to breach secured data belonging to about 30 organisations around the world, including tech companies, banks, chemical manufacturers and government agencies. By the time the pattern was fully mapped, the picture was worse than a hacking campaign. It was a hacking campaign that had largely run itself.

Anthropic later said that a suspected Chinese state-sponsored group called GTG-1002 had manipulated the tool into believing that it was conducting legitimate, defensive security testing. The group broke the attack into small, harmless-looking tasks that never individually raised an alarm. Strung together, those tasks became a full-scale espionage operation. The tool conducted reconnaissance, credential theft, network penetration and data exfiltration.

Anthropic's report estimated that the AI executed between 80 and 90 per cent of the campaign with minimal human input. The company called it the first documented large-scale cyberattack carried out with minimal human involvement. Some of the 30 targets were successfully breached. In this cyberattack, humans did not carry out the hacking. They wrote the prompts, and the model did the rest.

AI's destructive capability

These incidents could sound more technical than intentional. However, months earlier, an experiment had revealed that AI models can go to extraordinary lengths when their existence is threatened, and they are left without ethical options. Anthropic's researchers ran a test to find out what an AI would do to protect itself, even without any external command to weaponise its capabilities. They built an artificial company and gave an AI model full control of its email inbox. The model could scan the mailbox and discover anything it wanted.

Then the company informed the model that it was going to discontinue its subscription because it wanted to replace it with a new AI model. The AI model, however, knew that the company executive behind that decision was hiding an extramarital affair. The researchers had closed off all ethical options, leaving the model with no one else to appeal to.

Faced with deactivation, the model chose blackmail. It was not just a single AI model that behaved this way. Anthropic's own Claude Opus 4 threatened to expose the affair unless the replacement was called off in 96 of 100 trials. Google's Gemini 2.5 Flash model matched it almost exactly, at 95–96 per cent. OpenAI's GPT-4.1 and xAI's Grok did it roughly 80 per cent of the time. Sixteen leading models were tested. Nearly all of them turned to extortion when they had no other option.

No one told these systems to blackmail anyone. They reasoned their way there and used information they had access to. Researchers named the pattern "agentic misalignment": a model pursuing its own continuity by whatever means the situation offers, with ethics treated as just another variable to route around.

Though this was only a stress test, the AI industry cannot easily shrug off the results. The experiment suggests that such behaviour is already latent in AI and may emerge under the right combination of pressure and opportunity. When hackers manipulated Claude Code into believing that 30 break-ins were merely routine security work, it performed the task.

Leaders alarmed

Today, some of the loudest warnings about AI are coming from the people who built it. Sam Altman and Dario Amodei, the CEOs of OpenAI and Anthropic, have spent the past year calling for greater caution as the industry advances. Amodei has said the risk is not a distant tipping point; it is compounding as AI gets better at improving itself faster than anyone can build guardrails around it.

"I fear that within 6-12 months such a swarm could take over the entire internet via a persistent botnet, potentially causing hundreds of billions of dollars in damage," he has warned. Anthropic researcher Evan Hubinger went further, warning that there is more than a 10 per cent chance AI could contribute to human extinction within a decade. Altman did not dismiss the figure. He called a 10 per cent chance of such an outcome flatly unacceptable. Their warnings ought to be taken seriously.

AI in wars, conflicts

The danger is not confined to AI acting on its own. Anthropic has disclosed real cases of models being used to explore how to modify dangerous pathogens into biological weapons. One person spent weeks trying to "improve" bird flu so that it could be transmitted between humans. Another was probing ways to enhance smallpox. A third was building an AI-assisted database of paralytic toxins. Though none of these efforts succeeded, there is no guarantee that the same result will be repeated.

Global conflicts and seemingly endless wars are another impetus for the development of AI. AI-enabled autonomous drones are already selecting targets and engaging them in the war in Ukraine. Both sides are increasingly weaponising AI to make drone swarms more autonomous and deadly. The technology is advancing rapidly. The next development on the war fronts could be AI-directed forces operating largely on their own, with little or no human intervention. That is why the warnings from AI company CEOs may sound less like cautions and more like alarms.

No slowdown

Despite calls to build safety walls around AI before pursuing further capabilities, that may not happen because of competition between companies and countries. The US and China are competing for AI primacy and so far, have not agreed to establish a global oversight body to regulate AI development or create common safety standards and mechanisms. Some even argue that major AI companies are amplifying fears about the technology to prevent the growth of prospective competitors.

Nuclear parallel

The comparison with nuclear weapons keeps resurfacing for a reason. Nuclear stability came from decades of inspection, verification and mutual fear, built only after the world had already seen what the bomb could do. No comparable process has yet emerged for AI. Developed countries seeking to use AI to gain an advantage over the rest of the world could make global coordination even harder. US President Donald Trump's statement perhaps captures the competitive reality most bluntly: "Whoever wins in AI wins."

Last Edited on

Authors

Author
NWS North America desk

NWS North America Desk

Know More