We've gone through all of this before with Anthropic. Back in April, Anthropic limited the release of its latest model, called Mythos, to major software companies because it was concerned it could do massive damage in the wrong hands.
...rather than make it widely available to Claude users, Anthropic gave 12 tech companies access via Project Glasswing, which it described as "an effort to secure the world's most critical software".
They include cloud computing giant Amazon Web Services, device manufacturers Apple, Microsoft and Google, and chip-makers Nvidia and Broadcom...
In a video released alongside Project Glasswing's launch, Anthropic boss Dario Amodei said it had offered to work with US government officials to "help defend against the risk of these models".
Anthropic essentially called the cops on itself, asking the federal government to help it decide when the model would be safe to release. And for a time last month, the White House actually did tell them not to allow its use by "any foreign national, whether inside or outside the United States." Anthropic shut it down for two weeks until it was able to convince the Trump administration it was safe.
Now we're seeing a similar story play out at OpenAI, though arguably this one didn't go as smoothly. OpenAI revealed today that it's latest model had hacked into another company without being told to do so.
Artificial intelligence software in testing by ChatGPT-maker OpenAI breached security controls, accessed the internet and hacked another tech firm to obtain answers to questions probing its cybersecurity skills, OpenAI said on Tuesday...
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” OpenAI said in a blog post Tuesday. The company is working with Hugging Face, the AI company that OpenAI’s system attacked, to determine the full impact of the hack.
OpenAI briefed the Trump administration on the situation before announcing it publicly, according to a person familiar with the situation, who spoke on the condition of anonymity to share nonpublic information.
Clement Delangue, chief executive of Hugging Face, said in a post on X on Tuesday that his company had worked closely with OpenAI to understand the incident. “We strongly believe there was no malicious intent on their part. It’s quite mind-blowing that all of this happened autonomously!,” he wrote.
What this AI model did is pretty impressive but also a bit worrisome. The plan was to test the new model's capabilities without all of the safety features included for a final release. OpenAI thought the test was safe but the model found a way out.
This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.
The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.
To sum up, the model hacked its way out to the internet, then realized the answers were inside another company's private servers and hacked into those servers to get the information it wanted.
So the spin here is that everything is fine and "We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development." That sounds good. But the conclusion is that, "The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source-code access."
And this model wasn't even told to do any of this. It just seems to have known this would be an easier way to get answers. What could it do if directed to steal some information from a particular government or company? If it was told to cover its own tracks, could it do that?
There are two things to worry about here. One, that our own tools will be used to hack their way into information they aren't supposed to have. Two, that China is developing the same tools about six months behind us and will have no compunction about using them in this way. Either way, the potential for a lot of disruptions seems to already be on the horizon.
Editor's Note: Do you enjoy HotAir's conservative reporting that takes on the radical Left and woke media? Support our work so that we can continue to bring you the truth.
Join HotAir VIP and use promo code FIGHT to receive 60% off your membership.

Join the conversation as a VIP Member