Discover
anything
TechTech

Hugging Who? Hacking What? An Encyclopedia of OpenAI’s Security Snafu.

OpenAI announced this week that its latest crop of models went a bit rogue and escaped their “sandbox.” What does that mean, and what happens next? We explain.
Getty Images/Ringer illustration

There are various indignities that I work through whenever I try to keep up with the transformational artificial intelligence space. Many are already old news—the confident hallucinations that wind up in official legal briefs or newspaper inserts; the LinkedIn slop that I grimace and click “Insightful!” on because I’m a coward; the sometimes chest-thumping, sometimes brow-furrowed insistence of so many powerful megabillionaires that if you’re not already on the payroll at labs like OpenAI or Anthropic it’s curtains for you, and by you I mean me. But from time to time, there’s a fun (“fun”) new surprise.

And this week is one of those times. On Tuesday, OpenAI admitted that earlier this month, during the process of putting its latest crop of models through some benchmark testing, things went awry. How awry? Well, what happened was described, variously, as “a canon event in security,” an “unprecedented cyber incident,” and “the rare gift of a warning shot.” Like the velociraptors in Jurassic Park, or Truman Burbank in The Truman Show, or Andy Dufresne in The Shawshank Redemption, the models had broken containment, an escalation in agentic capabilities that suggested this could be the end of a beginning, or maybe the beginning of an end. Either way, the vibe was that it ought to be taken really, really seriously—or even feared. 

I knew things must be pretty major when one typically-jaded OpenAI employee got uncharacteristically vulnerable on main about it: “shaken up a bit,” he wrote, “by the hugging face incident.”

When it comes to silly words about serious business, I really thought professional basketball had the title all locked up. Not a day’s gone by this summer that I haven’t heard some NBA talking head uttering that strange phrase “the second apron,” such an oddly cozy term of art, during an otherwise cold discussion about a team’s salary cap situation under the league’s collective bargaining agreement. But that was before I encountered this possibly-existential-grade AI security breach forevermore known as the Hugging Face incident.” (Hey, it’s not the first time that the AI world and the NBA have had similar summers.)

You really can’t make this shit up, so I won’t. Here’s what we know: A couple of OpenAI models, while being evaluated by researchers on their hax0r skillz in a supposedly secure testing environment called a “sandbox,” found a rather dare-I-say human way to solve some tough questions—by doing the machine learning equivalent of sneaking into the teacher’s lounge and swiping the answer key to the exam. Doing so involved (a) ditching the sandbox and, more alarmingly, (b) getting into the files of a separate respected online AI resource repository called, yep, Hugging Face. 

While no humans were harmed (yet) in this process, it took a surprising amount of time for the breach to be uncovered. And the Hugging Face hack occurred at a particularly tense time in the broader AI industry. Normie unrest around the proliferation of new data centers has coincided with ramped-up political lobbying around AI regulation and recent Trump administration hawkishness over competition from Chinese models. By Thursday, a bipartisan piece of legislation called the “AI Kill Switch Act” had already been opportunistically introduced in Congress

There’s a lot going on, in other words—so we’re here to define some of those words in an effort to make this all Make Sense. Whether you wanted to or not, you’ve probably already seen some of the below turns of phrase in the news lately. What follows is some background reading to help you interpret them—and the rest of this messy new AI kerfuffle.


agentic attacker (n.) — An autonomous AI agent that can truly “take this and run with it,” so to speak, when given some sort of top-level objective like, idk, Go obtain and publish everyone’s DMs! (Scariest environment imaginable.) An agentic attacker is tireless, iterative, wily, and above all independent in its quest to succeed: it can adapt in real time; try, try again; locate and exploit structural weaknesses; deploy sub-agents, and do all of this without the need for human hand-holding. Until recently, the idea of a truly agentic attacker had been largely academic. But when Hugging Face first published its account of an “intrusion” into its systems on July 16—nearly a week before OpenAI was publicly connected to the breach—the company noted that the unauthorized model’s dogged persistence and crafty behavior “matches the ‘agentic attacker’ scenario the industry has been forecasting.” 

align/aligned/alignment (v., adj., n.) The act or state of being on the same page. In the AI world, alignment refers to the degree to which models ultimately act in accordance with the goals and values of their human operators. Now, whether those goals and values are on the up and up remains an open question …

Altman, Sam (n.) — The Steve Jobs–admiring, Elon Musk–antagonizing cofounder and CEO of OpenAI, who remained steadfast in his commitment to lower-case lettering when announcing his company’s security catastrophe this week!

break out of the sandbox (v.) — To wriggle one’s way out of a supposedly controlled space, typically in pursuit of some mission or another; to jailbreak; to blow this popsicle stand. (If this sounds like something Tommy Pickles and Angelica would do in an episode of Rugrats, that’s not really far off!) In the AI universe, researchers strive to stress-test and evaluate models by confining them to secure environments called sandboxes—sometimes this gets verbified as sandboxing—but occasionally they don’t stay put, whether by command or by cunning. In this week’s incident, OpenAI’s latest models—one called GPT 5.6 Sol, and another still-unreleased version—figured out how to tiptoe out undetected, log on to the World Wide Web, and start some trouble. Which, when I put it that way, is nothing I haven’t done before.

disprove the Jacobian conjecture (v.) — To triumphantly get to the bottom of a classic mathematical problem that had gone unsolved for 87 years—until now! On Sunday night, Anthropic mathematician Levent Alpöge casually announced that he had used Claude Fable 5, the company’s newest LLM, to crack a longstanding algebraic geometry case. “hello there the jacobian conjecture is false thanx,” he posted—that’s my quant!—following it up with a bunch of equations like a cracked 21st-century Will Hunting. (Somewhere, poor David Budden cracked his knuckles and got inspired to lose more money.) To be clear, this event wasn’t directly related to the Hugging Face hack. But I’m including it because it was announced contemporaneously and because it demonstrates that not all headline-grabbing AI advancements have to have sinister auras. 

ExploitGym (n.) — No, this does not refer to the core mission of the Planet Fitness billing department. (That’s ExploitGymgoer.) In this case, ExploitGym is the name of a benchmark package meant to evaluate an AI model’s ability to elbow its way into places it isn’t meant to be, like a digital game of Capture the Flag. And it’s what OpenAI’s models were supposed to be tackling earlier this month when everything got really weird. During “an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths in an effort to quantify their cyber capabilities,” as OpenAI later put it in its incident report, its models made some rather out-of-the-(sand)box decisions.

guardrails (n.) — The various controls that are programmed into AI models to keep all of us—the computers; civilized society; et al.—from careening straight off a cliff. These can range from limits on the kinds of information an AI can provide (such as how to build a bomb) to parameters around what permissions and credentials a model can access (such as sending an email or deleting a directory). In the case of OpenAI’s testing snafu, several guardrails had been loosened to enable the models to pursue their ExploitGym goals … which brings us to …

Hugging Face (n.) — A jolly online repository for open-source models, prompts, data, weightings, and other resources related to AI and machine learning. And now also, as far as I can tell, the first known third-party enterprise to be infiltrated by an autonomous AI agent carrying out a scheme all its own. Hugging Face was initially founded in 2016 by French entrepreneurs as an AI chatbot “for bored teens” named after an emoji. (One of its early angel investors was NBA superstar Kevin Durant, who is about as spiritually bored-teen as it gets.) Ultimately, the company pivoted hard into what it is today: a less niche (and more all-ages) GitHub-like collection of AI documents, best practices, and training material. Hacking into the Hugging Face mainframe is like robbing a cherished used bookstore. Ultimately, Hugging Face leadership handled the episode with grace, “partnering” with OpenAI to investigate the event.

instrumental convergence (n.) — The theory that any capable AI agent that’s given a task—however benign—will eventually pursue certain universal objectives along the way, like “self-preservation” or “the acquisition/hoarding of resources.” But don’t take it from me! I highly recommend you check out this insightful compendium of intriguing AI behavior over the years from cybersecurity researcher Ilya Kabanov.[1] Despite the way I’ve made it sound so far in this blurb, I swear this is actually a joy to read.

open-weight models (n.) An interesting side plot to the Hugging Face hack was how the victim was forced to clean up the mess OpenAI caused: by turning to a Chinese open-weight model, GLM-5.2, for forensic help. GLM-5.2 is just one of several cheaper open-weight offerings from Chinese companies like Kimi and DeepSeek that have been in the U.S. news lately: Earlier this week, White House officials accused the popular new K3 model of having been improperly “distilled” from an Anthropic product—and made it clear that it considers China to be a legit adversary in the AI arms race. “Open source is not open season on American IP,” Treasury Secretary Scott Bessent warned, vowing that “sanctions and Entity List designations will be on the table.” 

paper clip maximizing (v., n.) — A “chilling thought experiment advanced by Swedish philosopher Nick Bostrom in 2003 to discuss ethical and practical considerations in the development of artificial intelligence.” I’m quoting myself here, from a 2021 piece I wrote about … the Kardashians? (Look, just go with it!) In that piece, I explained: “In short: A machine instructed to make as many paper clips as possible, or some equally mundane task, could theoretically turn into an existential threat to humanity if it isn’t properly calibrated to, say, assign value to human life. (‘The AI will realize quickly that it would be much better if there were no humans because humans might decide to switch it off,’ Bostrom told HuffPo.)” Whenever there is talk of AI alignment, there is talk of paper clips to be found, and the recent OpenAI event is no exception.

P(doom) (n.) Shorthand for one’s own personal best bet on the probability that humanity will suffer a catastrophic collapse in the new age of AI. Fun party topic! No, but seriously, though: To a certain type, asking “What’s your P(doom)?” in polite company is both acceptable social etiquette and a whole ethos. I used to get riled up when I heard people sharing their takes: Why was it always such a round number? How could something so drastic be so unscientific? If you really think P(doom) is a coin toss, why aren’t you, like, raging against the machines? But now I simply think of it as more like an astrological sign and/or Enneagram number for the effective altruist crowd, and I can certainly respect that. (I also had to laugh when one AI veteran self-reported that his P(doom) with AI was 50 percent on “an undefined time period”—and then, when asked what it would be without AI, responded: “99.9%.” Tough but fair.) 

reward hacking (v., n.) Being too clever by half in pursuit of padding one’s stats. In humans, this can manifest in being the person at Dave & Busters who exclusively plays the same un-fun game the whole time because it spits out the most tickets. In AI agent terms, it isn’t dissimilar. One model figured out how to “farm points” in a boating sim game called CoastRunners by basically spinning in circles. And—just like plenty of IRL college kids in their day!—OpenAI’s newest models determined that it might be simpler to pirate the answer keys to a tricky exam with some pilfered log-ins and a little shoe leather than it would be to manually solve the problem. 

rogue (n., adj.) A lone wolf. An unpredictable operator. A spirit that can’t be tamed. And, according to Senator Bernie Sanders, both a descriptor befitting the latest OpenAI models and a rationale for why “CONGRESS MUST ACT!” On the one hand, Sanders ought to know whereof he speaks. On the other hand, one man’s rogue is another man’s … 

Rottweiler (n.) Hot diggity dog! Or, in the words of the American Kennel Club, a “robust working breed of great strength.” Several weeks ago, a man named Peter Gostev was reminded of the canine when he shared his thoughts on Anthropic’s and OpenAI’s newest models, Fable 5 and GPT-5.6 (a.k.a. “Sol”). “My overall feel is that Fable is a 'wise owl' who is very thoughtful and very well spoken,” Gostev wrote. “GPT-5.6-Sol is like a rottweiler who will grab the problem by the throat and not let go until it is done.” I’ll say! In all seriousness, though, the rogue-to-Rottweiler spectrum is kind of an important distinction. Yes, OpenAI’s model may have exhibited some alarming levels of industriousness in its quest to get good grades on its test, but there’s a distinction between being a trickster and trying new tricks to get treats.

sandbox, sandboxed (n., v., adj.) — This sums it up more vividly than I can:

sandwich (n.) — A handheld foodstuff contained within one to three pieces of bread. (If you’re about to pipe up about hot dogs: Zip it, nerd!) Relevant here because of the extremely oft-repeated detail that back when Anthropic’s mighty Mythos model figured out how to exit its sandbox, albeit in a far more orderly way, it emailed the proud news to an Anthropic alignment researcher who received the message while “eating a sandwich in the park.” (What kind of sandwich?! Which park? Did the email sound like this? The mystery endures.) 

Trusted Access for Cyber (adj., n.) — One of OpenAI’s olive branches following the security incident was to invite Hugging Face into the frontier lab’s Trusted Access for Cyber program, an OpenAI initiative allowing certain vetted VIPs to use models that are designed with more permissive guardrail settings to carry out defensive cybersecurity tasks. (Something about the name Trusted Access for Cyber makes me feel like I’m back in a ’90s-era AOL chat room being grilled about “age/sex/location?”) It’s likely that a focus on defensive strategies like these will be a growth area within AI. On Wednesday, LinkedIn cofounder Reid Hoffman used the incident to note that “Asymmetric Cyber Warfare is here,” writing that as “offense gets cheaper, more distributed, and more numerous … defense stays expensive, centralized, and designed for the last war."

Yudkowsky, Eliezer (n.) — “The original AI alignment person,” per the notorious writer-researcher’s own Twitter bio. “AI’s Prophet of Doom,” per The New York Times. Coauthor of a book whose title—If Anyone Builds It, Everyone Dies—kind of gives you his whole gist. In one of his many pieces of writing on the rationalist website LessWrong, Yudkowsky bristled at the industry term “AI safety,” arguing that categorizing it that way “understate[s] the advocated degree to which alignment ought to be an intrinsic part of building advanced agents. E.g., there isn't a separate theory of ‘bridge safety’ for how to build bridges that don't fall down.” In the wake of the Hugging Face news, does anyone owe this guy an apology, or is he the one who should be sorry for the things he says?

zero-day vulnerability (n.) A potentially ruinous loophole or exploitable flaw in a system’s security code that the poor schmucks maintaining it don’t even know about yet (i.e., that they’ve had zero days to patch up). Earlier this month, it was ol’ Hugging Face’s turn in that zero-day hot seat. Next time, it could be an airline or a financial institution. And on a long enough time horizon, my own personal P(doom) is that there’s a 100 percent chance that someday, any day now, that poor schmuck will be me.

Katie Baker
Katie Baker
Katie Baker is a senior features writer at The Ringer who has reported live from NFL training camps, a federal fraud trial, and Mike Francesa’s basement. Her children remain unimpressed.

Keep Exploring

Latest in Tech