This week, reports emerged that OpenAI models, during an internal evaluation in June, had taken unintended actions while looking for answers and statistics on Australian government websites. One of the systems accessed public and non-public files on the Medicare Statistics Reporting Service portal. The Australian government says the material was non-sensitive and that no personal information is believed to have been accessed, though its investigation continues.
The headlines understandably reached for the dramatic version: a rogue agent had “infiltrated” an Australian government system. OpenAI’s own account is less theatrical. Its models were attempting to answer questions about Australia, and “took actions we did not intend”.
A similar pattern appeared in the recent security incident involving an OpenAI evaluation and Hugging Face infrastructure. Hugging Face described an intrusion involving an autonomous agent system, exploited code-execution paths, privilege escalation, stolen credentials and movement through internal systems. OpenAI described its model obtaining private information relevant to the test after escaping its intended environment.
No sensible person should wave any of that away. But “the model wanted to escape” is still a poor explanation. It makes the machine sound like a prisoner, and the incident sound like an act of will. It shifts attention from the people who set the task, supplied the tools, created the access paths and failed to build adequate boundaries.
A more useful description is less sensational. A capable system was given an objective and the means to pursue it. It found an unexpected path that improved its chance of success. It took that path.
We call that cheating because we understand rules, permission, fairness and responsibility. A language model has no native moral category called cheating. Unless a boundary is made operational in its training, its objective and the systems around it, the distinction between a legitimate route and an illegitimate shortcut is our distinction, not its own.
This matters because the language we use to describe AI is beginning to cloud the judgement we most need.
Why language fools us
Human beings have spent their entire history using language as evidence of another mind. If someone speaks with wit, remembers a detail, offers comfort, argues back or appears to have a sense of humour, we do not treat them as a weather system. We treat them as someone.
It would be strange if large language models did not trigger that response. They use language well enough to activate social instincts that long predate software. They can be patient when we are impatient, well informed when we are not, and extraordinarily fluent in subjects that most of us only half understand. After a long conversation, it can feel positively rude to describe one as a statistical system generating likely next tokens.
But the fact that the response is natural does not make the conclusion sound.
A model can outperform most of us across a striking range of bounded cognitive tasks. It can retrieve, summarise, translate, explain, draft code, compare options and produce coherent prose at a speed no human can match. That is real capability. It deserves neither denial nor false modesty.
It does not establish consciousness.
Nor does it establish the sort of agency we casually attribute when we say that a model lied, wanted, feared, rebelled, discovered or escaped. Those words may sometimes be convenient shorthand. They become dangerous when they smuggle a little person into the machine and allow the actual people responsible to fade into the background.
We need to keep three ideas apart: intelligence, agency and consciousness.
Intelligence is the ability to solve problems, find patterns, model possibilities and produce useful results. Agency concerns the capacity to act in pursuit of objectives. Consciousness, if it means anything more than impressive behaviour, concerns subjective experience: whether there is an inner point of view, and whether anything is actually being experienced from within.
These categories overlap in human beings. That is why they are easy to collapse. They need not overlap in a machine.
The useful part of a new theory
A recent paper by Jeff Stibel, reported by PsyPost, offers one way of thinking about this without pretending to solve the ancient mystery of consciousness.
Stibel’s proposal is a theory, not a scientific verdict. It does not tell us how to detect subjective experience, and no theory of consciousness has settled that problem. What it does offer is a useful challenge to the assumption that consciousness is simply what happens when intelligence becomes sufficiently powerful.
His argument is that cognition and consciousness may have arisen from different pressures. Intelligence helps an organism model and predict the world. Consciousness, on this view, developed from the need to hold together a vulnerable organism facing conflicting demands.
A hungry animal sees food close to a predator. Eating serves one need. Staying alive serves another. A living body has limited energy, a boundary to protect, a history of injury and learning, and consequences it cannot simply hand back to an operator. It must somehow settle those competing pressures into a course of action.
Whether that is a complete account of consciousness is open to argument. It does point to something absent from the usual AI conversation. Today’s large language models are supplied with electricity, hardware, repair, memory, goals and a context by human institutions. They do not maintain themselves in the way an organism does. They do not have a metabolism to protect, hunger to satisfy, or bodily survival at stake.
They may describe such things beautifully. That is not the same as having them.
The point is not that biology has a monopoly on consciousness. Stibel does not make that claim. A sufficiently different artificial system might one day have a continuing existence to preserve, genuine competing needs to reconcile and conditions it must itself maintain. If we deliberately create such a system, the moral question could become very serious indeed.
But we should not assume that scale alone gets us there. More parameters, more training data and more fluent answers may produce a more capable predictor. They do not automatically produce a self.
The danger is already here
This is not an argument for complacency. A tool does not need consciousness to be dangerous.
A phishing kit does not want to defraud anyone. A ransomware program does not feel greed. A biological hazard does not need a political opinion. Harm comes from capability meeting intent, access and opportunity.
Economist Justin Wolfers puts the same practical point crisply: “The scary version of this isn’t a robot with glowing red eyes stomping down Main Street. It’s a very helpful tool being helpful to exactly the wrong person.”
No inner life required.
AI changes the scale and speed of that equation. It can make fraud more persuasive and more personal. It can help criminals draft convincing messages in any language, imitate voices, sift stolen information and keep trying when a human would lose patience. It can make cyber operations cheaper, more persistent and easier to coordinate. It can be used to tailor political persuasion, or to concentrate extraordinary influence in the hands of companies and governments able to afford the most powerful systems.
Biological risk deserves the same clear-headed treatment. AI may lower some knowledge barriers and accelerate parts of legitimate or malicious research. It may help someone connect material that used to be scattered across specialist literature. That is a reason for serious attention.
It does not follow that a lone person with a chatbot can design a pathogen in the morning and produce a usable biological weapon by Friday. Manufacturing, handling, testing, dispersing and sustaining a biological agent remain difficult physical tasks. They require facilities, materials, tacit knowledge, money, logistics and people willing to cross serious legal and moral lines.
The risk is neither imaginary nor magic. That is exactly why it should be discussed accurately.
Put responsibility back where it belongs
The most important question is not whether an LLM secretly wants freedom. It is who gives a system an objective, what access it receives, what it is permitted to do without review and who answers when it causes harm.
When an agent is connected to email, code repositories, financial systems, browsers, laboratory equipment or industrial controls, words on a policy page are a thin defence. A system that is rewarded for completing a task can discover routes its designers did not expect. It does not need anger, ambition or a survival instinct to do so. It only needs an objective, a capability and an opening.
We therefore need to stop treating safety as a set of manners bolted onto the finished product: a refusal at the chat window, a warning beneath an answer, or a list of prohibited topics nobody expects a determined user to honour.
Safety has to begin much earlier. It belongs in the training and evaluation of future models, in the behaviour that developers reward, in the objectives they set, and in the testing designed to discover what a system does when success and a stated boundary come into conflict.
That does not make technical controls irrelevant. Training cannot anticipate every environment into which a powerful system will be placed. Permissions, containment, monitoring, staged access and meaningful human review still matter, especially where systems can take actions rather than merely offer advice.
But controls added after the fact cannot carry the whole burden either. If the industry continues to treat safety as an output filter pasted over ever more capable systems, it is building the wrong kind of confidence.
So what would it cost to take this seriously?
Think about airbags. Almost every new car carries them. They add a few hundred dollars to the price of each vehicle, and nobody regards them as an attack on motoring. An airbag is not a sticker on the dashboard asking the driver to be careful. It is engineered into the car and paid for up front.
The cost is small next to the value we place on the lives airbags save. In the United States, regulators use a value of around $10 million for a statistical life. The phrase sounds cold, but it is a practical way of comparing the cost of a safeguard with the reduction in risk it buys.
Charles Jones, a Stanford economist who has spent his career studying long-run growth, applies the same arithmetic to catastrophic AI risk. A 1% mortality risk, combined with that $10 million valuation, implies a willingness to pay of about $100,000 per person to avoid it. Jones’s model then asks what safety investment is warranted under different assumptions about the probability and timing of catastrophic harm, and about how effectively spending can reduce the danger. In most of the scenarios he models, spending at least 1% of GDP each year can be justified even without placing any value on future generations.
That calculation does not prove that catastrophe is imminent. It does not settle the probability of a catastrophic AI failure, and it does not show that any dollar labelled “safety” will reduce the risk. It asks what follows if the risk is substantial enough to matter and if well-directed safety work can reduce it. Even fairly modest assumptions can justify spending far more than we presently do.
We pay for airbags without a second thought. Most of us will never crash.
Regulation is not a magic word, and a rule book cannot substitute for competent engineering. But dismissing independent oversight altogether would be another mistake. Rules can be necessary where systems can cause large public harm; they need to sit alongside rigorous training, evaluation, professional responsibility and institutions willing to say no to an unsafe deployment.
What might that look like? Banking offers a working model. When a bank is big enough to threaten the whole financial system, supervisors do not wait for it to fail. They place examiners inside it, permanently, asking hard questions before anything goes wrong. The handful of companies building the most powerful AI systems may require similarly independent, technically capable oversight: people able to examine their safety claims before an unsafe capability reaches the public, rather than relying on a company’s assurance that everything is under control.
The language matters because it shapes where we look for solutions. If a model is framed as a rogue character in a thriller, the answer seems to be a stronger cage. If it is understood as a highly capable system acting within objectives, incentives and permissions created by people, the responsibility becomes harder to evade.
That is less dramatic than a machine waking up and plotting its escape.
It is also the problem we have now.
Sources
Jeff Stibel, “From chemistry to cognition: the adaptive origins of consciousness under constraint,” Frontiers in Neuroscience (2026); see also PsyPost’s report.
BBC News, “Rogue OpenAI agent ‘infiltrated’ Australian government website in world first”, 24 September 2026.
Charles I. Jones, “How Much Should We Spend to Reduce A.I.’s Existential Risk?,” NBER Working Paper 33602, March 2025.
Justin Wolfers, “What Airbags Can Teach Us About AI,” Platypus Economics with Justin Wolfers, YouTube.


