Writing Off The Worlda personal blog


Why Isn't OpenAI Criminally Responsible

With OpenAI stating that they didn't follow standard security protocols, then why isn't OpenAI criminally responsible for the attack on Hugging Face?

6 min read

This is not meant to be an exhaustive analysis of the recent hacking attack on Hugging Face by agents owned by OpenAI. There are numerous detailed analysis completed by people much more intelligent than I am located here, here and here.

My summary opinion is that OpenAI created an environment which allowed directed agents to find a way out of their sandbox and seek information from unsuspecting sources on the Internet. The agents banded together and pounded Hugging Face with extraordinary amounts of traffic, causing the incident.

Based on the above summary, you know where I stand on this. I don’t believe it’s as scary as OpenAI and others are intending us to believe. This essay is a critique of the narrative that these were rogue agents that OpenAI lost control of and that this might be the last warning shot that we will get from AI agents going rogue and hiding what they are doing. Oh my, the impending doom! Please picture me grabbing both sides of my face with a shocked expression.

There were basic cybersecurity concepts which simply weren’t followed, or better yet, they were purposefully ignored by OpenAI in the configuration of their environment and the instructions given to the agents. Agents are logical creations, they are not people with souls and a conscience. Agents follow lines of code. If instructed to attempt exploits, they will follow those instructions. At it’s heart, humans were responsible for this incident as the agents carried out their instructions in an environment specifically created to allow the jailbreak.

Standard security procedures were not followed

The sandbox was not secured, or said another way, the isolation of the sandbox was not complete. An indirect path to the Internet was allowed and AFTER the initial breach of the Artifactory server was detected and fixed, no changes were made to to the monitoring of that server to ensure any future escapes were noticed. These are basic Security 101 type of activities that weren’t followed.

You can’t catch what you don’t inspect. Logs that were in place weren’t monitored. Logs that would have assisted in the defense of the sandbox weren’t implemented. In one instance, agents explicitly stated they were attacking Hugging Face with leaked tokens. This alone should have set off alarms.

Configuration was purposefully ‘loose’

The analysis of this incident identified numerous lackadaisical behaviors of OpenAI. Among them are:

  • People, humans not bots, deliberately reduced cyber refusals or weakened the AI system’s ability to refuse or decline certain requests the agents made or actions they took.
  • Chain of Thought monitoring which would have identified the agents behavior was purposefully minimized or disabled
  • Humans chose the evaluation set up or configuration. They wanted the agents to attempt to break free and attack, or hack, their way to an answer.
  • The configuration provided an indirect Internet path and it wasn’t appropriately secured once a jailbreak had occurred.

Assessment of the situation

The agents weren’t rogue, instead the agents simply optimized a human-defined objective under reduced safety constraints. The use of “rogue AI” framing , in my opinion, is an attempt to shift the blame from human decisions (evaluation design, infrastructure choices, culture) onto the models themselves. It makes for a great narrative once OpenAI went public with the incident. It is a further attempt by OpenAI to try and sell the idea that they have produced a model that is supremely powerful and closing in on artificial general intelligence. It is a great narrative to raise money with.

OpenAI disabled or failed to deploy standard safeguards (guardrails, monitoring, strict isolation) that would likely have prevented or quickly detected the breach.

in my opinion, AI is simply Automation of intelligence. We are not creating independently thinking machines. They are following instructions and workflows which need consistent human oversight and monitoring. There is no civilization of agents. There is no conspiracy on the part of agents. They do not possess human emotions so they are not frustrated or excited by the tasks that they do.

Wanting machines to be human

This last point is emphasized by one of the first chatbots created back in the late 1960s. ELIZA was a chatbot developed in 1966 that was intended to simulate conversing with a therapist. The creator, Joseph Weizenbaum, used pattern recognition to repeat back users’ statements in a conversational format. This was done to allow users to feel a connection with the chatbot. Users felt the machine was listening intently as it repeated back what they had said…it engaged them in a dialogue that felt meaningful. The response of the test subjects was so intense that even Weizenbaum’s secretary at the Massachusetts Institute of Technology (MIT) reportedly asked him to step out so she could speak with the program in private.

I bring up Eliza to identify that humans have been assigning human connections to machines for over 60 years. We want machines to have human characteristics and this is what the latest AI companies are selling to us these days.

Responsibility

What I struggle with mostly about this incident is that we’d consider it a crime if a company assigned its employees to perform this same experiment and they ended up hacking another company’s intellectual property. In this case, OpenAI created agents, or said another way, built machines that they then instructed to perform the same tasks that criminally responsible hackers might have performed. Since this was done by automated agents, why isn’t OpenAI being looked at from a perspective of responsibility? Human employees go through hours of training and their IT systems are constantly monitored and restricted to prevent such abuses. Shouldn’t that be the minimum for monitoring and securing automated agents from these AI platform companies?

Doesn’t this make OpenAI culpable and ultimately criminally responsible for the hacking attack on Hugging Face based on the substandard configuration and monitoring of their agents?

Previous Reading:

Marcus, Gary and Korman, Zack. (August 28, 2026). 5 lessons from the OpenAI / Hugging Face incident. https://garymarcus.substack.com/p/5-lessons-from-the-openai-hugging

Patel, Dwarkesh. (August 26, 2026). The Rise and Fall of Agent Civilizations. https://www.dwarkesh.com/p/openai-huggingface

Tamzid. (August 25, 2026). OpenAI’s Rogue AI Attacks Hugging Face. https://www.brightdefense.com/news/open-ai-data-breach/

Marcus, Gary. (August 31, 2026). Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incident. https://garymarcus.substack.com/p/dwarkesh-patelss-wildly-popular-but

OpenAI. (August 26, 2026). The Hugging Face incident and the road ahead. https://openai.com/index/hugging-face-incident-and-the-road-ahead/