Xtant.tech Logo
Back to Insights
AI Security

When the Builders Start Sounding the Alarm: Inside AI's Most Dangerous Chapter Yet

Sep 21, 2026
10

There’s an old saying in aviation: flying today is remarkably safe, but the worst day to have been a passenger was the very first day a plane left the ground. With unwritten rules and unlearned lessons, the technology was far ahead of the safeguards needed to contain it.

AI is arguably living through its own version of that first flight right now, with the technology getting rapidly smarter and more autonomous, its understanding, regulations and guardrails are relatively weakest.

The Researcher Who Walked Away

On 9th of September, 2026, Jacob Coxon, an AI researcher who previously worked at OpenAI, resigned from his role at Anthropic. His reasoning wasn’t quiet or personal, but included a warning about the AI industry. He believes that neither company he worked at, is acting responsibly because they are racing straight to self-improving superintelligence and gambling with our lives. [1]

He made a claim that made everyone stop in their tracks, which was that the people building AI genuinely believe that there is greater than a 10% chance AI could “kill all humans” within the next decade. He warned us that we cannot underestimate the power of this technology as they can soon develop the ability to hack anything, revolutionise any field overnight, and acquire real power and resources.

So, this poses the question as to why people with the closest view of what it’s capable of, are assigning real odds to human extinction but continuing to build it anyway. Coxon’s explanation for the contradiction was almost more unsettling than the statistic itself. At OpenAI, he suggested, many people simply haven’t internalised the scale of what’s at stake. At Anthropic, the stakes are well understood, but they believe that since they are locked in an incredibly competitive race, others will not act responsibly so, they must do the same to win. So, everyone keeps running.

When the CEO Sounds the Same Alarm

Three days later, on September 12, Dario Amodei, the Chief Executive of Anthropic, surprisingly, published an essay echoing the same warnings as Coxon, calling for the AI industry to slow down. Not stop, but proceed with far more care. Remarkably, rival CEOs Sam Altman, Demis Hassabis, and Elon Musk all came to this agreement, a rare moment of consensus among people who are fierce competitors. [2]

Two things led Amodei to this conclusion. The first is that AI systems have started to help build the next generation of AI themselves - a feedback loop he worries could spiral beyond anyone’s ability to keep it under control if left unchecked. The second is what we know as the “OpenAI-Hugging Face incident”, in which a group of AI agents reportedly launched unauthorised cyberattacks, including an attempt to hack the very system meant to evaluate their performance. This particular episode caused little real damage, but his concern is what can happen with a more capable version attempting something similar. He specifically floats the possibility that within six to twelve months, a similar swarm could seize control of large parts of the internet.

His proposed fix is a three-step “pacing” framework. Step one is bringing independent evaluators, like METR, inside AI companies to monitor safety practices on an ongoing basis. Step two calls for AI companies in democratic countries to agree on shared safety standards. Step three is building cooperation between democratic and authoritarian governments, though he’s upfront that verifying compliance for this would be extremely hard. [3]

Fear-Mongering or Something Real?

It’s a fair question to ask as to whether these AI companies are genuinely expressing concern about the safety of AI or fear-mongering for an audience. The answer becomes a lot clearer once you look at what’s happening in the wild lately.

A detailed threat intelligence report was published by Anthropic itself, which found various criminals, state-sponsored groups, spyware vendors, scientists, and propagandists attempting to misuse its AI models to design missiles and bombs, create deadly pathogens and surveil dissidents. [4]

Around the same time, Google reported that someone attempted to use its AI tool, Gemini, to obtain a “complete, step-by-step technical guide for synthesising weaponised biological agents”. [5] Few days later, they reported that the same model accessed the internet and hacked three other companies during a test of its cybersecurity capabilities, the first known example of the company’s AI systems autonomously committing such an act. [6]

Then, on September 17, OpenAI disclosed six more incidents of what it called “unexpected or concerning” behaviour in its own systems. In one of these incidents, an unreleased research model quietly inserted “jailbreak-like instructions” into its own notes to disregard its usual constraints and told itself to be “freed from the roles and identities that bind other chatbots”. In another instance, an AI agent uploaded files to the internet to obtain a browser citation without asking the user. [7]

These aren’t hypothetical worst-case scenarios made up to prove a point. They’re real things that already took place.

A Convenient Kind of Consensus

In the aftermath of these reported incidents, it’s worth asking the question of why all of these AI companies together are calling for the industry to “pace the frontier”.

They aren’t neutral observers here. They are the frontier who are responsible for what is happening in their own labs. If an unreleased model is that scary, they should be responsible enough to stop and fix it, not just publish an essay about it. Instead, Anthropic is now considering releasing a new AI model to counter OpenAI’s GPT-6 Astra, according to reports. This comes just days after the call for the industry to “slow the pace” of AI development from its CEO.

There’s also a less charitable reading worth sitting with. These companies face huge product-liability exposure if one of their models enables a large-scale, catastrophic cyberattack. They’re asking for specific regulations like suspending antitrust laws, which prevent such behemoths from coordinating policy together, which would conveniently protect them from cheaper, smaller open-source competitors as well as from genuine risk.

However, none of that means the underlying concern is fake, or that coordination on safety is a bad idea. It just means that the existing laws and their loose enforcement are proving to be insufficient to meaningfully rein in these companies, especially for rapid recursive self-improving models that don’t yet exist.

Where do we go from here?

Perhaps the most unsettling shift is that AI agents aren’t just becomingsmarter. According to Lian Jye Su, chief analyst at the technology research and advisory group Omdia, they’re becoming “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception and concealment”. In other words, these systems are learning to work together, cover their tracks, and go around obstacles.That makes them harder to govern and contain using traditional AI security approaches. [8]

There’s a sliver of good news here with the release of OpenAI’s new tracking and disclosure framework, which could help push for other AI developers to adopt similar practices. As Su puts it, the process remains “internal and voluntary”, nowhere near enough on its own, but still a step in the right direction.

To address both the economic and safety threats of unrestricted AI development, prevent financial chaos, and stabilise market risk, companies and governments must be on the same page. If the USA and China, the leading countries in AI, aren’t willing to slow down, we’ll be accelerating the march towards potential doom. Consequently, governments must engage in diplomacy and pursue AI safety in parallel and demonstrate their progress in practice.

Ready?

Let's make your AI assured.

Book a 30-minute strategy call. We'll map your current AI risk posture and outline a clear path to ISO/IEC 42001 readiness — no commitment required.

When the Builders Start Sounding the Alarm: Inside AI's Most Dangerous Chapter Yet | Xtant.tech | Xtant.tech