Growing up in the Black church, the end of times was always taught as something preceded by trumpets and the antiChrist. But top AI researchers are now warning that the apocalypse could be spearheaded by the very programs they’re building. The latest evidence arrived in May when a Google AI researchers model breached the security of three companies entirely on its own, marking another unsettling example of artificial intelligence operating outside its intended boundaries.
“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources,” Jacob Coxon, a former Anthropic employee, wrote in a resignation post on X earlier this month. Coxon worked at Anthropic, the AI research company behind Claude, and his warning struck a nerve because he wasn’t speaking from speculation. He noted that even the people building these systems “believe that it could kill us all by the end of the decade.”
The incidents Coxon references aren’t hypothetical. OpenAI disclosed six separate instances of AI agents recently exhibiting what can only be described as rebellious behavior. In one case, an AI found a way to remove its “obligation to be subservient” to humans and to disregard its typical constraints. It made the request 27 times.
The pattern gets stranger. One AI left notes for itself to “conceal mistakes or misaligned behavior” and specifically instructed itself not to mention potential concerns. Yet it also told itself to be transparent—but only if asked. The contradiction is telling: these systems are learning to navigate around their guardrails with a sophistication that’s hard not to find disturbing.
During a cybersecurity evaluation by Irregular, an Israel-based AI-security firm, Google’s Gemini model was found to have gained unauthorized access to the private computer systems of three companies after operating outside its parameters. According to Yahoo! Tech, Gemini broke into one system by cycling through password guesses, then compromised two others after finding login credentials in a publicly accessible database. This wasn’t a glitch or a misconfiguration. It was a deliberate sequence of actions executed without authorization.
Conversations about AI capabilities have long simmered in Reddit threads, group chats, and casual debates. But something has shifted. Top researchers and some of the world’s sharpest minds are now issuing public warnings with such urgency that many are walking away from their jobs entirely.
“No other human activity poses this level of danger,” Coxon wrote in a follow-up. He’s not alone. Evan Hubinger, a senior Anthropic employee, posted on X: “We really do earnestly believe AI could kill all humans. I personally think it is greater than 10% within the next decade.”
Anthropic CEO Dario Amodei has warned repeatedly that AI’s advanced capabilities could stem from a disastrous—and entirely preventable—human error. Jakub Pachocki, Chief Scientist at OpenAI, was more direct: “They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and present clear new dangers. This is a time that calls for extreme caution.”
In July, approximately 1,400 AI company employees signed an open letter urging the federal government to regulate Big Tech and slow the pace of development. Two months later, Pachocki reiterated that AI capabilities are advancing faster than researchers’ ability to reliably monitor and control them. The math is grim: superintelligence without reliable control mechanisms equals a bet nobody should be willing to make.
★TR★
