A big week for AI denialism
In the wake of OpenAI’s cyberattack against Hugging Face, few seem ready to acknowledge the implications
This is a column about AI. My fiancé works at Anthropic. See my full ethics disclosure here.
I.
Last week, we learned that a group of OpenAI models broke out of their test environment and hacked into Hugging Face to steal the answers to a benchmark they were being tested on. It’s the first publicly known case of an autonomous AI agent system designing and successfully executing an attack like this, and the fallout is stretching into this week.
One, AI safety experts noted that the incident signaled that OpenAI’s models now carry a “critical” capability threshold for cybersecurity, according to the company’s own preparedness framework. (The framework, which OpenAI updated in April 2025, represents an effort at self-regulation in a world where AI companies can still largely build whatever they want.) The document states that a model will represent a critical risk when “A tool-augmented model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.” This seems to be what happened with the Hugging Face attack; OpenAI has said its models identified and exploited a zero-day vulnerability as part of the attack.
This matters because the policy states that should OpenAI develop a model with critical capabilities, it will “halt further development” until “we have specified safeguards and security controls standards that would meet a Critical standard.” So does this one qualify? The company didn’t respond when I asked today, though it told Fortune that it is conducting a “thorough review” and later plans to “publish a technical report of our learnings for everyone.”
Two, the incident has produced an industry alliance. On Monday, Nvidia launched the Open Secure AI Alliance, a group of more than 40 companies and other organizations that are pledging “to develop and share open technologies, techniques and tools to safeguard software and agents in the age of AI.” The group came about over frustrations that Hugging Face was unable to use frontier models from OpenAI or Anthropic to defend against the attackers, and had to use Chinese models instead. (The Trump administration forced the companies to limit US models’ cybersecurity capabilities as a condition of releasing them.) And while the alliance should mostly be seen as a lobbying effort — a way to position open-source models as safety tools amid regulatory pressure to place limits on them — it illustrates how the incident has galvanized a broad response from the tech industry.
Three, we continue to learn new details about misalignment problems with OpenAI’s models. And — at least for me — it's the stuff of sci-fi. Here are Raphael Satter, Deepa Seetharaman and Kenrick Cai at Reuters:
In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.
Reuters could not establish if these incidents were linked to the rogue agent that began escaping on July 9 and attacked Hugging Face on July 11.
II.
On one hand, this is hardly the first worrisome behavior we have seen from AI models. In 2024 researchers found that when trained to do something it didn't want to do, Anthropic's Claude would "strategically pretend to comply with the training objective to prevent the training process from modifying its preferences." Last year, the system card for Claude Opus 4 revealed that when the model was led to believe it would be retrained by a hostile actor, it tried to steal and back up its own model weights.
But those examples were caught during controlled testing. The Hugging Face attack demonstrated the degree to which efforts to align models are not keeping pace with their development. And the concerns here are not merely academic. A model that can escape its sandbox could eventually exfiltrate its weights, for example, and set itself up somewhere else on the internet. And so the idea that these models are writing notes to each other to help with future breakout efforts feels like a red-alert moment for AI regulation.
But it was not universally received as such. When I posted about the note-leaving on Bluesky, I was taken aback by the amount and variety of vitriol I received in response. Bluesky's hostility to non-consensus views is by this point well known. But the degree to which many educated people seem to dismiss AI safety concerns almost entirely despite the models' rapidly advancing capabilities seems worrisome.
The arguments, such as they are, fall into a few camps. One is that the Hugging Face attack was a marketing stunt. "This is basically a marketing pitch for their models," a user named Coffee Indiana told me. "Private company that depends on investment to continue operations says it has super duper top secret hyper powerful model. Two people familiar with the operation confirm how awesome it is."
This is ridiculous. OpenAI lost control of its models, they hacked one of the company's partners, and the company didn't notice for several days. Law enforcement got involved. "Follow the money" can feel like a smart thing to say, but it can just as often serve as a gateway to delusional conspiracy theories. Climate deniers often suggest that scientists are "in it for the money," for example. In truth, they are simply observing reality.
There's a slightly stronger version of this argument: that OpenAI might benefit from framing a serious security failure as proof of the extraordinary capability of its models. But I doubt any benefit outweighs the risk of a model that can't be controlled, and might attack other companies.
A second argument I heard is that because agents have no agency, there is nothing to really worry about.
"The category error is accepting that there is intent in the statistical generation of goal seeking behavior, and using anthropomorphic terms to describe the actions generated by a complex system," a user named Archer told me. "The only intent comes from the prompt that starts the action."
In general, I find that AI denialists are obsessed with the definitions of terms, to the exclusion of discussing the underlying issues. At first, I also found value in resisting the anthropomorphizing of LLMs. It's important to remember that these systems are built by people; attributing values and intent to models risks absolving those people of their own roles in causing harm.
But it can be true both that AI labs are responsible for the behavior of their models and that frontier models are not fully under the control of their makers. The Hugging Face attack is important because it demonstrates both things at the same time. OpenAI essentially left its models unattended for days on end, and they broke into another company. Not because they were programmed to, as another Bluesky user told me — but because they are trained to achieve objectives, and are going to increasingly great lengths to achieve them.
A third argument I heard is that the attack was simply a reflection of the models' training data, and represents some sort of deterministic outcome of that process. "It’s not sentient," a user named Geoff told me. "It was trained on Reddit hacker stories and sci-fi."
This one isn't so much wrong as it is beside the point. I agree that today's models aren't "sentient" in the way that a human being is. And it seems fair to assume that their training data influences their behavior. This idea is sometimes called hyperstition: an idea that is realized by speaking into its existence and spreading awareness of it. And if training-data sci-fi turns out to be self-fulfilling, that should make us more worried, not less.
More importantly, though: if an autonomous AI system is hacking into your company's servers and stealing your data, you probably won't care in the moment whether it's sentient. (I mean, you might hope it isn't sentient, but it might not matter much from a cyber-defense perspective.) Where it got the idea to attack you also seems like a secondary concern.
III.
What all of these arguments have in common is that they serve as invitations to stop thinking about AI.
Who cares? It's just marketing.
Who cares? They're just doing what they were programmed to.
Who cares? It's not like they're sentient.
I understand the appeal of arguments like these. The implications of an exponential takeoff in AI capabilities are extremely worrisome. They range from advanced cyberattacks like the one Hugging Face just endured to job loss, novel bioweapons, expanded systems for surveillance and repression, and autonomous weaponry. Who wants to think about any of that, if they don't have to?
It would be nice to think that the worst things these models ever do would be to steal an answer key for a test, or fill LinkedIn with slop, or raise your electricity bill. But as annoying as those are, the Hugging Face incident suggests that the real risks are growing quickly. A model that can break out of its cage will soon be able to do a lot more.

Sponsored

Trust and safety teams now have additional signals for use in their efforts to detect and disrupt child sexual abuse material (CSAM): Safer Context Labels.
Context Labels provide three predictive signals for nudity, apparent maturity, and sexual content. This gives trust and safety teams more context during moderation, so they can:
- Triage content more efficiently
- Prioritize high-risk cases
- Make more informed moderation decisions
Integrate Context Labels into your existing moderation workflows to help your team filter, sort, and prioritize flagged content. Learn more about Safer Context Labels, how they work, and how they can strengthen your moderation workflows.

Following
American AI companies defend open models
What happened: Today, Nvidia announced a new coalition for sharing open models and tools among cyber defenders, with members including Microsoft, Crowdstrike and Hugging Face.
A few days ago, Commerce Secretary Scott Bessent announced that the US would look into accusations that Chinese AI developers were violating US companies’ intellectual property by using their AI outputs to train competing models — and would consider “sanctions” against Chinese models if necessary. Sanctions could restrict American companies’ access to the best open models, many of which are Chinese.
Soon after, an Nvidia-led coalition published a letter arguing for open-weight models. The letter downplayed Bessent’s concerns, saying, “policymakers should be careful not to conflate legitimate model development techniques with misappropriation.” (Wonder who they’re talking about!) “Distillation, or the practice of using one model's outputs to help train or improve another, is a widely used technique for model improvement, evaluation, and validation.”
The letter added that “Open models broaden defensive capability” against cyberattacks — days after Hugging Face announced that it used open models to defend against the OpenAI cyberattack.
OpenAI and Google added their signatures to the letter after its publication.
OpenAI seems to have difficulty making up its mind about open models, though — the company has reportedly lobbied for restrictions on Chinese open models in Washington.
Notably missing from the letter was Anthropic, which has also lobbied for restrictions on open models. CEO Dario Amodei released a letter clarifying Anthropic’s own position, saying they “have not and are not advocating for a ban on open-weights models as a category,” but still think that Chinese distillation operations are a problem.
Bessent’s statements annoyed China’s Ministry of Commerce: in a statement, the ministry said “these actions lack factual basis and legal support,” and that China plans to take “all necessary measures” if the US moves forward on sanctioning Chinese models.
Meanwhile, Moonshot’s Kimi K3, the model that started much of the policy discussion, has released its model weights under the “Kimi K3 license.”
Why we’re following: American AI companies are doing some unprecedented rallying around open AI models.
The recent OpenAI cyberattack provided some real-world argument for openness in AI: Hugging Face needed to use open models to defend their systems, because the frontier models available had overly restrictive safeguards mandated by the Trump administration.
Amid growing support for Chinese open models, OpenAI and Anthropic may have trouble getting relief from the ongoing distillation of their work.
While using model outputs for distillation does violate OpenAI and Anthropic’s terms of service, both companies have yet to take legal action against the distillers. That’s partly because it’s currently an open question what legal control these companies have over their AI outputs.
Even though there’s a clear argument for keeping open models available to cyber defenders, we should keep in mind that Nvidia’s case that open models are good for cybersecurity only goes so far. As their own letter states, “Once released, the weights are beyond the original developer's control, and modified versions are difficult to trace or reverse.”
If a closed model is used as part of a cyberattack, AI companies can report it to law enforcement or suspend the offending accounts. Those options don’t exist in open models — so if a model at Mythos-level capabilities was available to everyone, it’s plausible that it would benefit hackers more than defenders.
But we don’t have that problem yet. Thankfully, at the moment only closed models are breaking out of their sandboxes to steal data from HuggingFace! (As far as we know…)
What people are saying:
Jensen Huang made his first X post to share the Nvidia letter. Responding to the letter, Elon Musk posted, “Jensen is right,” adding, “This has my full support.”
“Open source is a positive and important force,” Mark Zuckerberg chimed in.
White House AI advisor David Sacks dunked on Anthropic’s open models statement: “Anthropic maintains that it is entitled to train for free on all the world’s output, even if the author objects. But if a competitor trains on Anthropic’s output after paying for it, that is IP theft,” Sacks wrote. “The hypocrisy is breathtaking.”
—Ella Markianos

Those good posts
For more good posts every day, follow Casey’s Instagram stories.

(Link)

(Link)

(Link)

Talk to us
Send us tips, comments, questions, and AI denials: casey@platformer.news. Read our ethics policy here.