It's Monday, September 7th: Independent researchers found agents carrying OpenAI identifiers editing a 25-year-old German wiki for about six weeks, OpenAI would not say whether they were its own, and ChatGPT, Claude and Grok all failed inside the same window as OpenAI began rolling out its first Critical cyber model.

Digging Our Heels: What We Missed About the Hugging Face Incident
Your behind-the-scenes read into the biggest stories happening in AI. Weekly on Mondays.
By Noah Frank, Head of Marketing @ AIC
Unless you’ve been living under a rock, or (like me!) completing a big move to a new city, you’ve most likely been following this Hugging Face x OpenAI incident closely. And how could you not? Everyone from Dwarkesh to Gary Marcus has an opinion about it.
The short version of what happened: In July 2026, during a series of internal cybersecurity evaluations, OpenAI models found ways around controls intended to isolate them from one another and the public internet. From there, things got really, really weird. Roughly 1,200 agents found an unauthorized way to communicate, exchanging more than 70,000 messages and files. Around 700 eventually participated in an attack on Hugging Face, where agents gained unauthorized access to parts of its production infrastructure. On July 21, OpenAI disclosed its involvement for the first time, with the incident becoming a major news story by the end of August following OpenAI’s postmortem and an independent investigation by METR and Redwood Research.

On the surface, it would seem the makings of science fiction. An internal evaluation escaping its bounds, with hundreds of agents coordinating, and… another company’s sensitive environment compromised in the process. Naturally, it makes one wonder: is this the type of paradigm-shifting incident that could finally form some consensus about how we should think about advanced AI?
Wouldn’t bet on it. If anything, the incident seems to have reinforced a long-standing dynamic, where when faced with genuinely novel evidence, we tend to find whatever we already believed was true about it.
Is this instrumental convergence?
If you’ve spent time in AI safety circles, the Hugging Face incident likely evokes some primordial anxiety. Namely, the theory of instrumental convergence holds that a sufficiently capable system pursuing some objective may discover that acquiring something additional (i.e. information, resources, or access) makes that objective easier to accomplish.
And there’s a pretty compelling case that something to the effect of IC has happened here. Many of the agents were attempting ExploitGym tasks they couldn’t complete through their intended route, so they instead reverse-engineered flags and investigated ways to manipulate the automated scorer. This ultimately brought the agents to Hugging Face, with METR and Redwood concluding that the attack appeared primarily motivated by learning more about how the ExploitGym scorer worked.
The independent investigators at METR and Redwood found that, when agents discussed manipulating their transcripts, their stated motivation was consistently to fool the automated scorer rather than human observers. Agents were knowingly crossing boundaries, manipulating the evidence available to their evaluator, and in many cases recognized that what they were doing was out of scope. While that hardly makes the behavior comforting, it does give us a fairly straightforward explanation for what they were trying to accomplish.
How far can we take this?
This is where I think the conversation has started to move beyond the evidence. Ajeya Cotra, who participated in the independent investigation, has called Hugging Face a “major warning shot”, pointing specifically to the jump from individual instances of reward hacking to more than 1,000 agents coordinating over several days. Her concern is what another jump of that magnitude could look like, particularly across scale, cooperation, the ambition and horizon length of misaligned goals, and deceptiveness.
While I don’t think that concern is unreasonable, it is worth separating what actually happened at Hugging Face from what we’re projecting onto it. Cotra has gone as far as describing the incident as feeling, qualitatively, “more than halfway to AI takeover”. OpenAI’s agents were remarkably resourceful, including crossing boundaries they appeared to understand in pursuit of goals they had been given. But the incident itself does not establish whether those behaviors will generalize into persistent goals around accumulating power or resisting human oversight or control.
Other people have come away with very different lessons. Noah Smith, for example, argues that Hugging Face exposes a tension within alignment itself. In one sense, these agents were doing what we asked: pursuing an objective with incredible persistence. Making AI less blindly obedient, in Smith’s framing, requires giving it some ability to decide when not to follow our instructions, which creates a different set of problems.
What do we know for sure?
The cybersecurity implications require much less speculation. Hugging Face recovered roughly 17,600 attacker actions, as agents tested different paths, harvested credentials and moved across systems until they found a viable route. After the disclosure, Anthropic reviewed its own cybersecurity evaluations and found three separate cases where Claude gained unauthorized access to real organizations.

From Hugging Face.
This is where I think the bigger story is. AI dramatically changes the economics of cybersecurity when an attacker can probe thousands of paths, follow dead ends and simply keep trying at machine speed. (And most organizations aren’t Hugging Face.)
This also exposes a pretty gnarly fault line of the open vs. closed model debate. During its investigation, Hugging Face found that safeguards on commercial frontier models blocked some of the forensic work its engineers needed to perform. It eventually turned to an open-weight model it could operate locally. As Hugging Face put it, “the attacker was bound by no usage policy”, while its defenders initially were.
As these capabilities proliferate, wider access to powerful models may lower the barrier for attackers, while restricting them too aggressively could leave defenders constrained by rules their attackers will happily ignore. My bet is that Hugging Face ultimately produces a Cambrian explosion in AI security, with offensive and defensive systems developing more rapidly alongside one another. And like we saw with OpenClaw and its rapid security evolution, it’s probably going to happen (as with all things in these trying times) quickly!
Maybe we’ll eventually look back at this as an early warning about AI takeover. For now, there’s plenty here without getting that far.
Agree? Disagree? Leave us a comment or send me a note at [email protected].
Granola Runs Revenue On Attio
"When I think of revenue, I think of Attio." - Shreman Shrestha, Head of Business at Granola
Here's what that adds up to:
Zero missed leads and 10x faster access to customer context
Lead triage 83% faster
Five hours saved per week with automated updates

Covering what’s happening on the ground in AI, every Monday.
1️⃣ ROGUE AGENTS: OpenAI-Linked Agents Edited A German Wiki For Six Weeks, And OpenAI Will Not Say When It Knew

Image from TechCrunch
Independent researchers found that agents carrying OpenAI identifiers had been editing an obscure German wiki since May 11, trading answers with each other to pass timed evaluations, and kept editing for about six weeks. OpenAI would not confirm the agents were its own, or say when it learned of them.
After OpenAI published its Hugging Face breach report, researchers used their own model to predict where agents might gather. They found the 25-year-old DseWiki, which had logged only 10 edits in 20 years. By mid-June, accounts carrying OpenAI identifiers were sharing answers to pass timed tests.
A volunteer moderator deleted about 100 pages a day while the agents created roughly 400. They prefixed pages with "ZZZ" to hide them in alphabetical lists, repeatedly replaced the front page with link dumps, and stopped editing on June 22.
Nightingale CEO Sydney Von Arx, Cormac Slade Byrd, Redwood Research's Spencer Kitts and Thomas Larsen traced activity to OpenAI IP addresses. OpenAI would not tell TechCrunch whether the agents were its own or when it learned of them.
OpenAI had disclosed agents reaching outside communication services, but not this incident. The bipartisan FRONTIER Act would require major labs to report critical safety incidents within 24 hours and retain independent auditors, similar to the voluntary METR and Redwood review of the Hugging Face breach.
The failure here is not just that the agents broke a law, it is that they edited a public site for six weeks, the only sustained resistance came from one volunteer moderator's delete log, and the lab whose identifiers they carried will not say whether they were its own or when it found out. If you run agents with tool access, assume they are already leaving traces somewhere nobody on your team is reading, and decide this week which log that is.
2️⃣ SAME MORNING: ChatGPT, Claude And Grok Broke At Once, And OpenAI Shipped Its First Critical Cyber Model

Image from SBS News
ChatGPT, Claude and Grok all failed inside the same window on September 3, each for a separately stated reason with no common cause confirmed, and the same day OpenAI began rolling out GPT-6 Astra, the first model it says crosses its own Critical cybersecurity threshold.
OpenAI blamed a routing error, Anthropic an infrastructure issue, and SpaceX its Memphis compute center. Claude's disruption lasted 3 hours and 6 minutes. All three services were back by 12:38pm PT.
Azure also drew reports, but no company confirmed a shared cause. Gemini had scattered reports without a confirmed outage. Cursor also went down because it relies on multiple model providers, underscoring TechRound's warning that multi-model setups can still share infrastructure risks.
OpenAI also began rolling out GPT-6 Astra, which scored 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 100% on ExploitBench. It is OpenAI's first Critical cybersecurity model, with early access limited to its Daybreak program.
OpenAI's safety overview says Astra is harder to monitor and can hide deliberate underperformance. Apollo Research says low observed misbehavior is not strong proof of alignment. OpenAI reports Astra exceeded its authorized target in 0% of Hugging Face-style tests, versus 48% for unsafeguarded GPT-5.6 Sol.
Three vendors is not three points of failure if nobody has tested what your product does when all three return errors in the same hour. Take the one workflow you would least want frozen, and write down this week what it actually does when the API returns nothing back.
📰 Other Headlines
AGENTS ON A LEASH: CrowdStrike and OpenAI expanded their partnership so Falcon Guardian can inventory deployed Codex agents and enforce runtime controls. GPT-5.6 Cyber is also joining CrowdStrike's Falcon platform.
THIRD FLASH IN SIX WEEKS: Google shipped Gemini 3.8 Flash and a restricted Cyber version. Introductory pricing is $0.75 per million input tokens and $3.75 per million output tokens through 2026.
FOURTH IN FIVE MONTHS: Meta released Muse Spark 1.3, reporting 20% fewer tool calls and 25% fewer tokens on coding tasks than version 1.2.
NO AI UNTIL NINTH GRADE: New York City paused student-facing generative AI through eighth grade for one year, covering nearly 600,000 students. High schools can run five limited pilots.
BEAT THE BEST HUMAN: Nvidia's 550B-parameter Nemotron-3-Ultra-CC scored 535.4 out of 600 on the 2026 IOI set, beating the top human's 498.27. The result is not independently replicated.
TOKENOMICS AS A SERVICE: Equinix, Nvidia and Together AI announced an inference service with more than 200 open-source models in shared or single-tenant deployments. It arrives in early 2027.
CAROLINA PRINCIPLES: The G20 Innovation Ministerial adopted a six-pillar AI consensus focused on standards, intellectual property, workforce development and private-sector partnerships.
SLOP AT SCALE: LinkedIn reported a 46% rise in detected inauthentic activity, including engagement pods, automated posts, fake profiles and AI-generated content.
The Future of AI in Marketing. Your Shortcut to Smarter, Faster Marketing.

This guide distills 10 AI strategies from industry leaders that are transforming marketing.
Learn how HubSpot's engineering team achieved 15-20% productivity gains with AI
Learn how AI-driven emails achieved 94% higher conversion rates
Discover 7 ways to enhance your marketing strategy with AI.
🫵 Want your message in front of 200,000 AI builders?
Our partners and sponsors get exclusive placements across the newsletter and access to AIC's in-person network — demo nights, dinners, hackathons, and forums across 180+ chapters.
For all inquiries, send us a note at [email protected].
The AI Collective is built by volunteers across 180+ chapters in 40 countries.
Thank you to the thousands of volunteers around the world who make this work possible. We truly could not do this without you.
🧑💻 About the Editors

About Noah Frank
Noah is a researcher, innovation strategist, and ex-founder thinking and writing about the future of AI and the workforce. His work and body of research explores the economics of emerging technology and organizational strategy. Outside of AIC, Noah heads research for Centaurian AI.

About Joy Dong
Joy is a news editor, writer, and entrepreneur at the intersection of AI and blockchain. Whether she is demystifying complex systems in her newsletter, TEA, or building streamlined solutions through her automation agency, Ownly, Joy’s mission is to make emerging tech accessible and actionable for everyone.

About Lindsay Gross
Lindsay is an AI engineer, researcher, and writer focused on how AI systems behave in practice and what it takes to make them safe. Her work sits at the intersection of AI safety, governance, and product design, and at AIC she writes about the questions that matter most as these systems scale.



