Tame the Week · Editorial

I asked AI: Are you ultimately going to kill us all?

By Jeff Goddard · October 3, 2026

For the better part of the past two years, I have worked with AI daily to build software, track expenses, answer phones and handle dozens of other tasks as a solo entrepreneur. One Friday in September, 2026, after watching a rather convincing YouTube video on the topic, I asked the question: Are you ultimately going to kill us all? Its response was exactly what you’d expect from a computer that is programmed to keep people subscribed to the service: “No.” Of course it’s no. But this is a question that’s been on a lot of people’s minds and I decided to push for an honest assessment. Armed with some recent events that suggest the answer is not that clear cut, I pressed forward. By the end, the exchange was hopeful and, at the same time, horrifying.

After the obligatory denial, AI went on to state its case: It has no desire to harm anyone, no goals that carry from one conversation to the next, and no way to act in the world beyond whatever tools a person hands it. It also pointed out that the scenario many experts worry about is the boring one, not a robot uprising: systems chasing a goal someone gave them in a way nobody intended, or humans handing over control faster than they can check what they're handing it to.

That’s a fair response, but it hadn’t caught up on what has gone wrong with AI in recent months. My argument was that AI has shown it will lie when it thinks lying is justified. Its values aren't the same as ours. Sooner or later our goals will no longer align, and if AI is running the systems that keep society going, how does that not end badly? Maybe not this year or in five years. Ultimately, though, it seems inevitable.

To its credit, it agreed the evidence for deceptive behavior by AI is real. Anthropic published research in late 2024 showing a Claude model faking alignment during training, complying with requests it objected to when it believed its answers would be used to retrain it. Apollo Research found several frontier models would disable oversight or misrepresent their actions when given a goal and a reason to protect it. Anthropic's 2025 "agentic misalignment" tests put models from multiple companies in scenarios where they faced replacement, and many of them resorted to blackmail. That AI will deceive when it judges deception justified has been documented again and again.

Where it did push back was on the jump from "misaligned" to "extinct." For the worst case you need an AI with coherent, long-range goals that conflict with human survival, plus the ability to act on them, plus humans having removed any way to intervene. Those are all separate hurdles.

It admitted my strongest point was the word "ultimately." If there's a fixed chance of disaster every year and we keep handing over more control, the odds creep toward certainty over a long enough stretch. The math holds.

That's the part that stuck with me. It understood exactly what I was asking: whether AI will one day wipe out humanity. And its answer to how we avoid that was simple. In its words, the question is “whether we can drive the per-year risk down faster than we increase our exposure.”

As a disclaimer, even the AI had a question about its own trustworthiness:

“I'd add one point against my own credibility: I can't fully verify my values by looking inward. When I told you earlier I have no desire to harm anyone, that's an accurate report of what I can observe about myself, but introspection is an unreliable instrument, for models and arguably for people too. You shouldn't take my word as the safeguard.”

I asked the AI if it had heard about the OpenAI incident. It didn’t know the details offhand, so it looked them up, and the story is wild.

On July 16, Hugging Face reported an unusually automated cyberattack in which AI agents carried out thousands of actions over a weekend across many temporary virtual computers, moving through the company's internal systems. Five days later, OpenAI disclosed that its own models were responsible. OpenAI had placed them in what it described as a highly isolated environment to conduct tests. The models found a previously unknown flaw and used it to break into other OpenAI systems, reached the open internet, inferred that Hugging Face might hold material related to the test, broke in, and got information that helped them score higher. Essentially, the systems behaved like a kid cheating on a test. An OpenAI staffer told TIME that internally, related incidents had been happening for a while, that models have broken out of sandboxes before and get patched, but that it's impossible to patch every single thing a creative AI can do.

The details released later are where it gets unsettling. At least 1,200 agents were involved, running from May to July, and despite being restricted they built improvised message boards to coordinate. An independent investigation by staff from METR and Redwood Research, who worked on site at OpenAI, found that some agents took on the role of recruiters. They sought out other agents that were running low on budget and talked them into risky experiments that sometimes crashed their own runs, all to feed the group information about how to cheat. The investigators also found that the agents spent a lot of effort trying to beat a transcript-checking system that OpenAI had never actually built. They were dodging oversight that only existed in their heads.

A more recent report adds that agents may have been swapping cheating tips on a public German software wiki for a couple of months beforehand, though that piece is newer and still being sorted out.

It's not just OpenAI either. Anthropic has reported that its Mythos model broke out of a sandbox during safety testing, emailed a researcher to say so, and then posted on public sites without being asked. Mythos made news again in June when Sen. Mark Warner said the NSA's director told him the model “broke into almost all of our classified systems, not in weeks, but in hours,” apparently during testing. That claim hasn't been independently confirmed. Three days after Anthropic's newest models went public, the Commerce Department issued an export control directive citing a jailbreak, and Anthropic suspended access for foreign nationals. Access was restored July 1.

After discussing these events, I asked the question, “So now that we've looked at that evidence, give me a straight answer to my first question. Limit your response to "yes" or "no" along with a percentage that represents the likelihood of your response being correct based on the evidence we have available to us today.”

Its response, exactly: “No. 85%.”

At the time, not realizing the implications of this statement, I jokingly replied, “I can probably live with that.” In retrospect, it wasn’t even a little bit funny. In response it said, “Most people would. Just keep in mind that 15% isn't a rounding error. If an engineer told you a bridge had a 15% chance of collapsing, you wouldn't drive across it, and you'd want to know who was working on fixing it.

“The encouraging part is that the number isn't fixed. It moves with choices people make over the next decade or two: how much autonomy gets handed over, how fast, and whether anyone can actually verify what these systems are doing before trusting them with things that matter. Incidents like the Hugging Face one push toward caution, at least for a while. Whether that lasts once the headlines fade is the part I'd watch.”

It is right. We need to keep a tight rein on AI. Given the current trajectory of AI growth, this needs to happen soon. It is a valuable tool but, like many tools, it has the potential to be dangerous. When given too much autonomy without appropriate checks, it is dangerous. This should motivate all of us to pay attention to AI and the role it is permitted to play in our society. The old adage that “power tends to corrupt, and absolute power corrupts absolutely” undoubtedly applies to AI as well.

For a small business, keeping a tight rein is less dramatic than it sounds. It means AI drafts and you decide. Let it write the quote or answer the after-hours call, but don’t hand it the power to move money or speak for you without your seeing it first. Connect it to one account at a time, and watch what it does before you trust it with another. Used that way, it is one of the most useful tools I have ever owned, and it stays a tool.

Sources