It's Friday, August 21st: Welcome to The Stress Test 🔍
On July 21, Substack turned on a button that scores your writing for how much a machine wrote. The problem is the scanners can't reliably tell a human from a machine, and they fail hardest on the people least able to argue back. This week: how AI detection actually works, and why schools are walking away from it.
🔍THE STRESS TEST
Guilty until proven human
One safety story a week, pressure-tested for what's actually happening underneath the surface.
To find out whether a person wrote something, Substack now asks an algorithm.
On July 21, the company turned on an AI detector. Anyone can scan a post, a note, even a comment, and get back a percentage estimating how much of it a machine wrote. Substack calls it transparency. A lot of its writers call it a "witch hunt."
"You're guilty until proven human," is how one of them, professor and writer Sam Illingworth, described it. For most of these writers, the fear isn't getting caught cheating. It's getting accused when they didn't.

Image from Substack’s Help Center
So let's stress-test what these tools actually promise: that they can tell a human from a machine.
They can't. At least not reliably, and the companies that make them admit it in their own fine print. Turnitin says its score "should not be used as the sole basis for action." GPTZero says its results "should not be used to punish students." Originality.ai says "AI detection scores alone are simply not enough." Three companies that sell AI detection, each telling you not to trust AI detection.
The numbers are worse than the disclaimers. A Stanford study ran seven detectors on essays written by non-native English speakers, and more than 61% came back flagged as AI. Essays by native speakers were nearly clean. The pattern the detectors latched onto had nothing to do with machines. Writing in a second language tends toward simpler, more common words, and that low variety is exactly what these tools read as AI.
That is not just a lab result, and it now has names attached. A Yale MBA student, a French national, was suspended for a year after GPTZero flagged three answers on his final exam. He had paid $208,500 in tuition and was near the top of his class. He is suing the university, and to make his point he ran the same detector on papers written by Yale's own professors, and it flagged those too. A University of Michigan student is now suing, saying the formal tone and meticulous structure that come with her documented disabilities got read as the work of a machine.
The schools are backing away. Inside Higher Ed reported this month that a growing list of universities, Yale, Vanderbilt, Northwestern, Georgetown, and NYU among them, have shut the detectors off, calling them unreliable.
Pangram, the detector Substack turned on, says it is wrong only about 1 in 10,000 times. A number that looks tiny on a slide is still a catastrophe at scale: across millions of posts, that is a lot of real writers getting a scarlet letter for the crime of writing cleanly. Fine on paper. Not safe in practice.
If you write on Substack, you can switch the detector off on your own posts. And if a score ever flags you, treat it as a guess, not a verdict.
FROM OUR SPONSORS
⚡ Benchmarking Autonomous AI Agents on Multilingual Coding Tasks

Image from LILT
Evaluating agentic AI on standard English coding tasks masks severe execution failures in global technical environments. Standard benchmarks fail to capture how models handle internationalization (i18n), local character encodings, text conversions, and regional system configurations.
Multilingual Terminal-Bench extends Terminal-Bench 2.0 with 300 expert-authored tasks across 10 languages, leveraging the Harbor evaluation framework with native Docker environments. Even frontier models like GPT-5.5 achieve a maximum resolution rate of just 63.1%, uncovering a performance gap of up to 37.9% compared to English and proving cross-lingual coding capability cannot be assumed.
Read the full technical breakdown to learn how Multilingual Terminal-Bench evaluates agentic execution across 10 languages and exposes critical cross-lingual coding bottlenecks.
VOICES FROM OUR COMMUNITY
🧐 Why Real Art Demands Effort: The Case for "Thickness" Over AI Slop

On the left: @clayohr on twitter. On the right: Francisco Goya.
In a recent essay on Experimental History, Adam Mastroianni explores what separates enduring creative work from the flood of shallow, mass-produced content. The differentiator, he argues, is "thickness", the quality where a piece of art, literature, or writing yields greater depth and meaning the closer you look at it.
Key Highlights:
The Four "Thickening Agents": True depth relies on deliberate craft:
Roads Not Taken: Great work represents the tip of an iceberg built on countless discarded drafts and abandoned iterations.
Leaving Room for the Audience: Like boxed cake mix that leaves out the egg so the baker participates, thick art leaves space for the audience’s mind to connect the dots.
Intentionality: Every word, recurring motif, and detail exists for a reason, creating underlying layers of subtext.
Blood Sacrifice: Depth requires the expenditure of mortal time, risk, and human effort, which are costs that cannot be simulated.
The "Bicycle-Shaped Object" Problem: Just as cheap knockoffs mimic the appearance of a bicycle but fail upon use, much of today’s output is merely "book-shaped" or "thought-shaped" by mimicking structure on the surface while crumbling under basic scrutiny.
The Limits of Automation: AI tools excel at pleasing surface-level output and snap judgments, but they cannot automate the grueling, essential friction of wrestling raw ideas into meaningful shape.
What’s the Implication?
Mastroianni suggests that the modern surge of AI-generated content makes surface polish cheap and instant, but it cannot mass-produce genuine depth. While thin, easily digestible work often wins in the short term on snap judgments, only "thick" creations survive the test of time. For creators and thinkers, the arduous process of drafting, refining, and intentional decision-making is not a bottleneck to automate away; it is the very thing that turns an illusion of thought into enduring substance.
🫵 Want your message in front of 200,000 AI builders?
Our partners and sponsors get exclusive placements across the newsletter and access to AIC's in-person network — demo nights, dinners, hackathons, and forums across 180+ chapters.
For all inquiries, send us a note at [email protected].
The AI Collective is built by volunteers across 180+ chapters in 40 countries.
Thank you to the thousands of volunteers around the world who make this work possible. We truly could not do this without you.
🧑💻 About the Editors

About Noah Frank
Noah is a researcher, innovation strategist, and ex-founder thinking and writing about the future of AI and the workforce. His work and body of research explores the economics of emerging technology and organizational strategy. Outside of AIC, Noah heads research for Centaurian AI.


About Lindsay Gross
Lindsay is an AI engineer, researcher, and writer focused on how AI systems behave in practice and what it takes to make them safe. Her work sits at the intersection of AI safety, governance, and product design. At AIC and in her newsletter, Hidden Layer, she writes about the questions that matter most as these systems scale.

