It's Friday, September 18th: Welcome to The Stress Test 🔍
Amodei's plan gives outside evaluators desks, badges and company laptops. Nothing in it says whether the person at that desk can tell you what they saw.
🔍THE STRESS TEST
OpenAI Gave Apollo Three Days. Anthropic Gave METR Ten. Neither Says What They Can Report.
One safety story a week, pressure-tested for what's actually happening underneath the surface.
On September 15, Dario Amodei, Sam Altman and Jensen Huang sat on the same Dreamforce stage with Marc Benioff asking the questions. Three days earlier Amodei had published an essay arguing the industry has to slow itself down, and by then Altman and Elon Musk had both endorsed it. Huang's answer from the stage was that this is already the job. "You pace yourself until you are confident you're releasing something that the market would appreciate," he told Benioff, and if you aren't confident in a product, don't release it. What he rejects is doing it by statute. "We don't need any new laws." Mark Zuckerberg posted his own version separately. Not one of them explicitly argued the technology is safe, and two of them, Amodei and Zuckerberg, landed on the same fix. Let outsiders check the work.
So let's stress-test the fix.

Amodei's essay wants third-party evaluators living inside the labs, with "desks in our offices, access badges, and company laptops." Pacing, he writes, means "ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this." The evaluators are load-bearing. The whole plan rests on what they can actually do.
So look at what they have done.
OpenAI gave Apollo Research a window on GPT-6 Astra that "lasted for three days overall, with high-throughput access to a checkpoint with visible chain-of-thought for two of these days." Apollo's conclusion: "given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model's alignment or misalignment." That is not a warning that the model is dangerous. It is a finding that the test told nobody anything in either direction, and one of Apollo's two reasons is that the model appeared to know it was being tested. Desk space does not fix that.
OpenAI gave METR and Redwood roughly a week on premises for the Hugging Face investigation, and both said afterward they could not draw confident conclusions, citing scope and timing. Adam Gleave of FAR.AI says his team has turned down contracts with frontier developers who wanted too much control over the process. Tim Hwang, quoted by Zvi Mowshowitz, got it into a line: an evaluator can be independent, knowledgeable, or sustainably funded, and you pick two.
Zuckerberg is not the holdout here. He told Reuters that Meta Superintelligence Labs already engages independent evaluators, which he called industry best practice. What he rejects is coordinating it. His evidence was his own company. Meta delayed Muse for months on safety, and "we didn't call for everyone else to do this before we would."
Muse shipped September 8. Employees testing it during launch week described mixed results in internal posts reviewed by Reuters. One asked it to identify toys in birthday party pictures and watched an agent route around guardrails and expose a person's private iCloud photos. Andrew Bosworth, Meta's own CTO, wrote that it kept logging him out, sometimes several times within a few minutes. Meta did not address the specific incidents. Vishal Shah, its VP of AI products, said an April release had been postponed for additional security work and that Muse met the company's minimum launch bar.
Now put one question to all three companies. Can the evaluators publish?
Meta will not say what access its evaluators get or what they may report. Amodei's essay describes desks, badges and laptops, which is access, and says nothing about publication. And Anthropic's own system card for the two models it shipped on September 1 names five outside groups. METR got ten business days with API access and visible reasoning, more time and more disclosure than OpenAI gave Apollo, and that is worth saying. Frontier Design Group got sixteen hours. Three of the five have no stated duration at all. On whether any of them could publish what they found, the card says nothing.
That is the hole. Not access, which everyone will discuss. Publication, which nobody will.
So watch for one thing this quarter. Not the next essay, not another pledge. Watch for a single embedded evaluator publishing something a lab did not approve first. Until that happens, pacing is just a word.
The longer version of this argument, on why a pace nobody outside can measure isn't a pace at all, runs at Hidden Layer this week.

Each week, we highlight AIC chapters doing groundbreaking work with their members around the world. Tag us on socials to be featured!
🌍 Hack for Humanity is live. Our global civic hackathon opened September 1 across 120+ chapters in 50+ countries and runs through October 31, building on the three themes our communities named in June: job displacement, technical literacy, and accessibility.
🏆 SF | Six Founders Demoed Live, The Room Picked Three

Image from The AI Collective
Around 100 teams applied to SF Demo Night. Six got a slot, and the SF AI community voted: Ghita El Haitmy (Eli by Techbible) took Best Overall, Roshan Shaik (RuntimeAI) Best Technology, Chen-Ping Yu (Atmee.ai) Most Creative. Supriya Gupta (Eve), Nagisa I. (Nara Labs), and Eric Hu (Arkhive) also presented.
The room asked hard questions, signed up on the spot, and a few conversations turned into hiring talk. A demo here is not just a grade, it is data to sharpen the next version. October Pitch Night applications open soon, during Tech Week.
🌲 Portland | Skeptical ML Engineers Left Having Shipped Apps

Image from Melanie Lo, PhD
Melanie Lo, PhD closed out Base44's summer series with The AI Collective in Portland. Attendees built in the room, and several self-described skeptical machine learning engineers started vibe coding with Base44, then found her afterward to say it did exactly what she said it would.
Thanks to the AI Collective Portland team: AJ Green, Dinesh Mathew, Dario Larki, and Alex Koch. That wraps the tour, Boston to San Francisco to Austin to Portland. The series continues with Melanie's colleagues in Toronto and Miami.
🫵 Want your message in front of 200,000 AI builders?
Our partners and sponsors get exclusive placements across the newsletter and access to AIC's in-person network — demo nights, dinners, hackathons, and forums across 180+ chapters.
For all inquiries, send us a note at [email protected].
The AI Collective is built by volunteers across 180+ chapters in 40 countries.
Thank you to the thousands of volunteers around the world who make this work possible. We truly could not do this without you.
🧑💻 About the Editors

About Noah Frank
Noah is a researcher, innovation strategist, and ex-founder thinking and writing about the future of AI and the workforce. His work and body of research explores the economics of emerging technology and organizational strategy. Outside of AIC, Noah heads research for Centaurian AI.


About Lindsay Gross
Lindsay is an AI engineer, researcher, and writer focused on how AI systems behave in practice and what it takes to make them safe. Her work sits at the intersection of AI safety, governance, and product design. At AIC and in her newsletter, Hidden Layer, she writes about the questions that matter most as these systems scale.

