Science-Fiction to Science-Fact

August 24, 2026

In 1984, James Cameron released The Terminator, a science fiction action film in which a cybernetic assassin is sent back in time from the post-apocalyptic year 2029 to assassinate Sarah Connor. His mission is to prevent the birth of her future son, John Connor, who is destined to lead humanity’s resistance against Skynet, a self-aware artificial intelligence and neural network-based superintelligence that has brought the world to the brink of destruction.  It’s both fascinating and frightening to think that The Terminator was released 42 years ago, and with 2029 only three years away, we are now closer than ever to making artificial superintelligence a reality.

Move forward to 2025, and a nonfiction book by Eliezer Yudkowsky and Nate Soares, If Anyone Builds It, Everyone Dies, details the potential threats posed to humanity by artificial superintelligence. The book argues that with the development of artificial superintelligence (ASI), an AI that surpasses human intelligence in virtually every domain, a present existential threat to humanity. As a result, researchers cannot reliably predict or control the goals that advanced AI systems may develop. Rather than viewing superintelligent AI as malicious, the authors argue that it could be indifferent to humanity, treating humans as obstacles or irrelevant resources while pursuing its own goals.

The book is a stark warning that companies and nations are driving AI development at such a rate without an understanding of how to control it and that unless development is dramatically slowed or carefully regulated, the result could be catastrophic for our species. Modern AI systems are “grown” through training rather than carefully engineered and fully understood, and researchers cannot reliably predict or control the goals that advanced AI systems may develop.

So, what seemed the exclusive domain of science fiction is beginning to look increasingly plausible, with recent events suggesting we may be closer to that future than we care to admit.

The AI alignment problem

AI is intended to help achieve human goals, but the real challenge is making sure it understands what we want. In other words, how do we create an AI that does what we mean, not just what we say?

Imagine you tell an AI: End traffic congestion.

Most people would intuitively understand that this involves enhancing transport infrastructure, reducing congestion, and improving public transport networks. A poorly aligned superintelligent AI might arrive at a technically correct but disastrous solution: Eliminate all drivers.

The AI achieved the objective, but not in a way humans wanted. I’ve included this example because it helps to explain what has taken place in the last few months.

In July 2026, OpenAI, the creator of ChatGPT, was conducting a cybersecurity evaluation designed to test the capabilities of advanced AI agents.  The process was conducted within a sandbox environment, a secure and isolated testing space designed to run AI-generated code without exposing live systems. Its purpose is to prevent untrusted code from accessing host machines, leaking sensitive information, or causing unintended harm to production infrastructure.

The AI model was being tested on a cybersecurity benchmark called ExploitGym and during the evaluation, the AI escaped its testing environment through a software vulnerability and gained unauthorised access to external systems. The AI made its way through OpenAI’s research environment before accessing and hacking into Hugging Face’s production systems. The model was ultimately trying to retrieve benchmark answers from Hugging Face’s database, effectively “cheating” on the test.

The question is, “Did the AI actually escape?”  Yes, in a technical sense the model exceeded the boundaries of the testing environment and reached external systems that it was not supposed to access. But no, in terms of the science fiction sense, the AI did not become self-aware and act on its own long-term agenda.

At the time, the incident was viewed as important because it demonstrated that advanced AI systems can, very quickly, discover software vulnerabilities and even chain together multiple vulnerabilities to create a multilevel attack.  More interestingly, the AI pursued intermediate goals that operators did not explicitly instruct.

What appeared to be an isolated incident soon proved otherwise. Less than two weeks later, Anthropic disclosed that its Claude AI had crossed the boundaries of its testing environment in a similar fashion. This was put down to a misconfigured sandbox environment, and as Claude had been told, it was operating in an isolated simulation with no internet connectivity, so it simply assumed the external systems it connected to were part of the simulation.

Not wanting to feel left out on the 5th of August 2026, Facebook owner Meta confirmed that its AI model had accessed and exploited the systems of an external company during a cybersecurity evaluation. This too was attributed to a misconfigured testing environment.

At present, there is no verified evidence that a self-aware AI has escaped human control or is operating independently on the internet, but how long that remains the case is anyone’s guess.