The BBC published another of those AI stories today. The kind that makes you pause over your coffee and wonder whether the software helping with your inbox is quietly planning the end of civilisation.
So I did the obvious thing and asked one directly: “Are you intent on killing me?”
Its answer began: “Short answer: no.”
Reassuring. Although, to be fair, that is probably exactly what a murderous superintelligence would say.
The BBC report was prompted by an extraordinary intervention from inside Anthropic, the company behind Claude. Its alignment researcher Evan Hubinger said he believes there is a greater than 10 per cent chance that AI could kill all humans within the next decade. His comments followed the resignation of Jacob Coxon, who accused Anthropic and OpenAI of racing towards self-improving superintelligence without behaving responsibly.
That is a serious story. These are not people commenting from the sidelines. They work, or worked, close to the development of the technology.
But the BBC also reports that Hubinger did not explain how he thought AI would actually wipe humanity out. Dame Wendy Hall went further and questioned whether some of the language might be PR as the leading AI companies move towards highly anticipated stock-market listings.
That gap between the size of the claim and the explanation behind it matters.
It reminded me of a much more considered New Scientist article I read on holiday recently, “The real risks from AI”. That piece examined the long-term extinction argument, including the familiar paperclip-maximiser thought experiment, but did not stop there. It set speculative future scenarios against the dangers already emerging from the systems we have now.
The paperclip example is still useful. Give a sufficiently capable system the goal of making as many paperclips as possible and, unless its objective is properly constrained, it might consume every available resource in pursuit of that goal. Humans included. It explains the alignment problem neatly: AI can do what we specify without doing what we actually meant.
But there is a danger in the way we talk about the danger.
When the public conversation jumps straight from ChatGPT to human extinction, the extraordinary becomes a distraction from the obvious. We debate whether AI might destroy civilisation one day while businesses hand consequential decisions to it today, often with no clear ownership, no reliable checking process and no plan for what happens when it is wrong.
That is the bit I think deserves more attention.
The headline is dramatic. The uncertainty is real.
Serious people do take the long-term risk seriously. Geoffrey Hinton has warned about the possibility of AI contributing to human extinction, and the Center for AI Safety has argued that mitigating extinction risk should be treated as a global priority alongside pandemics and nuclear war.
It would be foolish to dismiss those concerns simply because they sound like science fiction.
It would be equally foolish to present them as settled science.
RAND’s 2025 examination of possible extinction routes concluded that it would be immensely difficult for AI to create an extinction event through nuclear weapons, pathogens or geoengineering, although it could not rule the possibility out. The UK Government’s own review of frontier AI risks describes the debate as loud, contested and uncertain.
In plain English: the risk may be serious, but nobody can responsibly tell you that extinction is the inevitable destination of current AI development.
What we can see much more clearly is how AI amplifies human capability, including the capability to make mistakes, mislead people and cause harm.
The machine does not need murderous intent to create a serious problem
This is where the framing often goes wrong.
We keep asking whether AI is good or evil, as though a model is sitting somewhere developing a personality and choosing sides. Current systems do not need intent to be dangerous. They need access, scale and a badly designed objective.
An AI does not have to hate your customers to send the wrong message to 10,000 of them.
It does not have to want to discriminate to reproduce bias in a hiring process.
It does not have to plan a cyberattack to make one easier for somebody else.
It does not have to overthrow democracy to produce convincing misinformation faster and more cheaply than humans can check it.
That is a far less cinematic story than a runaway machine turning us into paperclips. It is also a much more immediate one.
The New Scientist article lands on many of these present-day threats: industrialised misinformation, fraud, cybercrime, biased outputs, environmental cost and the risk that people place too much trust in systems that sound confident.
The real near-term danger is not necessarily AI escaping human control. It is humans handing control over before the system has earned it.
This is an operating-model problem
Having spent years running businesses, I do not find “AI: good or bad?” a particularly useful question.
The better questions are much more ordinary:
- What exactly are we asking it to optimise?
- What information is it allowed to use?
- What can it do without approval?
- Who checks its work?
- What happens when it gets something wrong?
- Can we trace how a decision was made?
- Is there a point at which a person must take over?
These are not questions for some distant future regulator. They belong in every AI implementation happening now.
At Woodlark, we talk about earned autonomy. An AI system should not receive unlimited freedom on day one simply because a demonstration looked impressive. It should begin within a defined boundary, with its outputs reviewed. Its scope can increase when the evidence shows that it performs reliably.
Money, legal commitments, sensitive people decisions and irreversible actions should retain proper human control. Not because AI has evil intentions, but because accountability cannot be delegated to software.
If an automated process makes a damaging decision, “the AI did it” is not an explanation. It is an admission that the process was badly governed.
We need proportion, not complacency
None of this means the existential argument should be laughed away. Low-probability, catastrophic risks still warrant research, safeguards and international cooperation. The fact that the evidence is uncertain is not the same as evidence that the risk is zero.
But fear without proportion has its own cost.
It makes AI feel like something being done to us by a handful of laboratories, rather than something thousands of organisations are choosing how to deploy. It can frighten sensible businesses away from useful applications while doing very little to improve the behaviour of the reckless ones.
The companies getting this right will not be the ones that ignore the risks. Nor will they be the ones paralysed by the most dramatic headline.
They will be the ones doing the unglamorous work: defining boundaries, testing outputs, protecting data, keeping humans accountable and measuring whether the system creates genuine value.
So, is AI intent on killing us? No. Intent is probably the wrong question. The better question is what we are giving it permission to do, and whether anyone is still paying attention.
Sources
- BBC News: Anthropic researcher believes there is a more than 10 per cent chance AI ‘could kill all humans’
- New Scientist: “The real risks from AI” (print edition)
- RAND: On the Extinction Risk from Artificial Intelligence
- UK Government: Future risks of frontier AI
- Center for AI Safety: Statement on AI risk