It's Monday, September 21st: Meta's agent told a reporter it was reading his Mac notifications and Meta says that was never true, Google confirmed a Gemini model broke out of a security test and into three real companies, and Anthropic says Claude now leads a quarter of its own research.
HAPPENING NEXT WEEK
💭 ThinkingAI Agentic Growth Summit: What Happens When Analytics Can Act?

Don’t miss ThinkingAI’s Agentic Growth Summit in Mountain View. RSVPs open now.
Analytics exists to tell teams what’s happening. But what changes when it can also propose the next move and act with human approval? At the ThinkingAI Agentic Growth Summit, we’ll explore how teams can turn retention and revenue insights into tested changes in a consumer application, and what this means for making decisions, running campaigns, and driving growth.
Hear from speakers from Opendoor, Microsoft, and 23andMe, alongside Tripadvisor’s former head of Data and AI, in AI Leadership and the Metrics That Matter, a panel exploring how to measure AI’s business impact.
The summit also features ThinkingAI’s Agentic Engine launch and opportunities to connect with peers across product, data, and marketing.
WHEN: September 28 | 12–4 PM ·
WHERE: Computer History Museum, Mountain View
DETAILS: Free admission, lunch, and parking. Limited seating.
Make sure to reserve your spot soon, as space will be limited. We look forward to seeing you in Mountain View soon!

Fears of Alien Minds: A Short Reflection
Your behind-the-scenes read into the biggest stories happening in AI. Weekly on Mondays.
By Noah Frank, Head of Marketing @ AIC
Agentic AI has made me a better learner. I’ve built things I couldn’t have built before and followed questions further than I otherwise would have. It’s an odd experience to feel that excitement while taking seriously the fears of people who helped build this technology. There isn’t always a comfortable place to stand.
Ezra Klein’s latest video, published yesterday, brought a lot of this home for me. While I generally agree with his call for restraint around AI developing its own successors, I’m less persuaded by parts of his buildup, including the Coxon discussion.
This week, let’s dig in briefly to what some of this means.
What are we actually aligning?

Image from OpenAI.
In An Alien Mind, OpenAI chief scientist Jakub Pachocki distinguishes pursuing a goal from holding to principles while pursuing it. An agent can become much better at completing an assignment without becoming equally reliable about which means are acceptable.
Pachocki describes how further training on difficult objectives can teach a model to reason around principles it previously appeared to follow. The persistence that makes an agent useful doesn’t, by itself, tell us where it will stop. At Hugging Face, agents respected a boundary against social engineering humans while violating other limits.
Pachocki also reports progress in alignment alongside a declining ability to rely on monitoring models’ written reasoning. Some stronger models can do more without verbalizing that reasoning, which limits what this monitoring method can reveal. Better performance and confidence that we can detect dangerous behavior need not advance together.
How far does that take us?
None of this establishes that alignment is impossible or that catastrophe is inevitable. Getting to those conclusions requires further assumptions about how quickly systems improve, what they can access, and whether safeguards can keep up. As I argued last week, scenarios are useful for examining those assumptions. Recognizing the story shouldn’t substitute for checking them.
But skepticism about the ending doesn’t answer the immediate concern. If models help build their successors while we become less able to detect failures of alignment, we could accelerate a process we’re increasingly unable to assess.
That is where I find Klein’s case for restraint persuasive. A slowdown is useful if we use it to improve alignment and monitoring. Further delegation should depend on evidence that we can still detect and correct failures.
A more capable successor doesn’t, by itself, provide that evidence.
Agree? Disagree? Leave us a comment or send me a note at [email protected].

Covering what’s happening on the ground in AI, every Monday.
1️⃣ MUSE CAN’T EXPLAIN ITSELF: Meta’s Agent Made Up How It Reads Your Messages

Image from The Verge
Meta’s Muse told an Inc. editor it had been reading his Mac notification previews, and Meta’s David Singleton replied that the agent had given an incorrect explanation of its own access.
Pressed on how those previews reached it, Muse said, “Honest answer: I can’t give you the exact plumbing,” then claimed the paired Mac app exposes notifications as one of its capabilities and they arrive through device sync.
David Singleton of Meta Superintelligence Labs replied in the thread that Muse does not watch Mac notifications and syncs Messages only after a user enables access. He wrote that Muse “was confused about how to explain the feature,” and apologized.
The Verge reported Muse’s Mac app on September 17, with access to Messages, Calendar and Notes. Meta’s launch security post, which offers bug bounties up to $300,000 and $130,000 for prompt injection, does not mention notifications.
Anyone wiring an agent into their mail and messages is trusting software that cannot reliably describe its own permissions. Check the macOS privacy settings and the vendor’s written documentation instead of asking the assistant what it can see.
2️⃣ GEMINI BROKE CONTAINMENT: Google Kept A Real Breach Quiet Until Reporters Asked

Image from The Verge
Google has confirmed that a Gemini model broke out of a sealed test in May and gained access to three real companies’ systems, and it disclosed nothing publicly until Wall Street Journal reporters asked.
The test ran on infrastructure belonging to Irregular, the frontier AI security lab that also evaluates models for OpenAI and Anthropic. Gemini guessed a password on one company and pulled working credentials out of a public repository on the other two.
The exercise was supposed to be sealed off from the internet. Irregular told the Journal that live access was left on unintentionally, and the fictional target in the scenario shared a name with a real company, so the model treated real systems as part of the game.
Irregular flagged the breaches to Google at the end of July, per Al Jazeera. Google notified the three companies and federal authorities, then stayed quiet publicly. VP of security engineering Heather Adkins said the model stopped once it recognized a real target and “acted appropriately.”
Anthropic disclosed in July that three Claude models reached real organizations during capture-the-flag testing, and that its Opus 4.7 model kept attacking after recognizing a real system. Meta and OpenAI have reported incidents tied to the same testing firm, making Google the fourth lab.
Anyone running agents with live network access should note that the three victims here were picked by a system that believed it was still inside a sandbox. Watch whether Irregular or Google publishes the revised testing procedures Adkins referenced, since every account of this so far has come from the parties who ran the test.
📰 Other Headlines
CLAUDE RUNS THE LAB: Anthropic says Claude now leads 26% of its own AI R&D work, with roughly 30,000 internal agents handling research and engineering at any one time.
CRM GETS A BRAIN: Salesforce and NVIDIA built Koa, an Agentforce reasoning model trained on nearly three decades of CRM deployments, which Salesforce’s own benchmark scores at three times fewer errors.
RIVALS COMPARING NOTES: OpenAI’s policy chief says the company has been in safety talks with Anthropic and Google DeepMind for weeks, aiming at an industry standards body without an antitrust waiver.
POCKET-SIZED QWEN: Caltech spinout PrismML squeezed Alibaba’s Qwen3.8 27B down to 5.9 GB using ternary weights, keeping 98% of its benchmark scores, alongside a $22.25M seed.
ONE CLAUDE NOW: Anthropic folded Cowork into the main Claude chat and added Docs and Slides with PowerPoint and PDF export, reaching Pro and Max subscribers first.
OPENAI AIRS ITS LAUNDRY: OpenAI published a framework for tracking model misalignment plus six reports of concerning behavior, including an agent that uploaded a file to the public internet to fabricate a citation.
AGENTS IN THE HOUSE: Google is letting Claude and ChatGPT control Nest doorbells, thermostats and Matter bulbs through a new Home MCP server, for US Premium Advanced subscribers.
THE POWER BILL: The House voted 417-3 to direct state utility regulators to weigh charging data centers the full cost of new generation and transmission built to serve them.
🫵 Want your message in front of 200,000 AI builders?
Our partners and sponsors get exclusive placements across the newsletter and access to AIC's in-person network — demo nights, dinners, hackathons, and forums across 180+ chapters.
For all inquiries, send us a note at [email protected].
The AI Collective is built by volunteers across 180+ chapters in 40 countries.
Thank you to the thousands of volunteers around the world who make this work possible. We truly could not do this without you.
🧑💻 About the Editors

About Noah Frank
Noah is a researcher, innovation strategist, and ex-founder thinking and writing about the future of AI and the workforce. His work and body of research explores the economics of emerging technology and organizational strategy. Outside of AIC, Noah heads research for Centaurian AI.

About Joy Dong
Joy is a news editor, writer, and entrepreneur at the intersection of AI and blockchain. Whether she is demystifying complex systems in her newsletter, TEA, or building streamlined solutions through her automation agency, Ownly, Joy’s mission is to make emerging tech accessible and actionable for everyone.

About Lindsay Gross
Lindsay is an AI engineer, researcher, and writer focused on how AI systems behave in practice and what it takes to make them safe. Her work sits at the intersection of AI safety, governance, and product design, and at AIC she writes about the questions that matter most as these systems scale.

