Part four of four, and the one the other three were building toward. The series starts here.
On September 2, a developer typed an ambiguous instruction into a terminal and the machine stopped.
Not an error. Not a refusal. A question. Meta had released Muse Spark 1.3 that morning, and the model’s new habit is to ask what you meant when your prompt is unclear, to ask for help when it gets stuck, and to confirm before it does anything consequential. Meta’s launch post leads with this, before any benchmark.
The most interesting thing about a frontier model, in September 2026, is its manners.
Sit with how strange that is
Clarify on ambiguity, escalate on failure, confirm before destruction. This is Norman. This is Nielsen. This is the second and fifth of the ten heuristics. It is the “are you sure?” dialog, which has been a product-layer decision since the Macintosh.
For forty years, whether a system interrupts you, and how, and at what threshold, was ours to choose. It lived in the design file.
Now it ships in the weights, at $1.25 per million input tokens, and it arrives with a personality you did not author.
That relocation has a consequence founders should think about before their next architecture review. The interaction pattern is becoming a procurement decision. Pick Muse Spark 1.3 and your product hesitates. Pick a model tuned for autonomy and it does not. You will inherit a set of judgements about when your users deserve to be asked, made by a lab optimising for its own benchmark suite, and you will discover them in your support tickets.
The old joke was that engineering decisions become UX decisions whether or not a designer is present. The new version is that model selection is now an interaction design decision, and it is usually made in a room with no designer in it.
The timing is what makes it uncomfortable
On the same day, Margaret Mitchell, Avijit Ghosh and Samir Passi posted the revised version of a paper called “AI Agents Push Humans Out of the Loop.” Its argument runs directly at the confirmation dialog.
Human oversight, they write, is the industry’s favourite answer to agent risk, and it is not a simple one: current agent designs impede effective oversight, and the cognitive capacities oversight requires are themselves eroded by extended use of the systems being overseen. Automation research has known this for decades under the name skill atrophy. The paper’s closing warning is that agent systems, as built, “passively incentivize the degradation of the very human skills they rely on.”
So one lab shipped the approval button, and the same week the research asked whether anyone is still qualified to press it.
The third fact, and the one nobody is pricing
Artificial Analysis measured Muse Spark 1.3 reading about 57% more input tokens per task than its predecessor, at $0.55 per task against $0.40. Two of its evaluation scores went down, which the firm attributes to a higher abstention rate. The model declined to answer more often when it was unsure.
In production that is exactly what you want. On a leaderboard it reads as a regression.
We have built an entire measurement apparatus that scores confidence and cannot see restraint. Meanwhile OpenAI’s GPT-6 Astra hit 99.9% on ARC-AGI-3 the following day, on its own harness, which means that particular test has finished its useful life. When capability saturates, behaviour becomes the product. Behaviour is design.
And this is where NN/g’s Anna Kaley and Raluca Budiu supply the sentence that ties the whole series together. Writing about what they call the custodial era of UX, they observe that “production has become cheaper than UX evaluation. It can now take less time to create an experience than to decide whether that experience is useful, usable, trustworthy, or coherent.”
Three independent sources, one shape. Generation is collapsing in cost. Judgment is not. The bottleneck in every AI product built next year will be the speed and quality of human decisions about output, and almost nobody is designing for that.
We are treating this as a governance problem
It is a design problem wearing a compliance costume.
The instinct in most companies is to answer agent risk with an approval workflow, a policy document and a log. All necessary, none sufficient, because every one of them assumes a competent reviewer at the end of the chain and none of them produces one.
Approval is a surface. If it does not carry the diff, the blast radius and the undo path in a single view, it does not transfer understanding, it transfers liability. And a reviewer who approves five hundred agent actions a day at 400 milliseconds each has not been kept in the loop. They have been given a stamp and a reason to stop reading.
We know how this ends because we already built it once. Cookie consent began as a transparency rule and became the most reflexively dismissed surface on the internet. Article 50 of the EU AI Act came into force on August 2 with real teeth, up to €15 million or 3% of global turnover, and the disclosure it mandates is now a design brief. We can make it meaningful or we can make it another banner.
Judgment cannot be shipped as a feature
After twenty years of teaching designers, the one thing I am certain of is that nobody develops judgment by reviewing finished work. They develop it by choosing badly and living with the result.
Which is precisely the experience an approval queue is engineered to remove.
If the interface only ever hands people a plausible answer and a confirm button, we are running a training programme in reflexive assent, and the Mitchell paper is right that the cost compounds quietly. Somebody has to design the friction that keeps a person genuinely deciding, and that somebody is us.
To be fair to the other side, a model that asks too much is a model nobody keeps. Confirmation fatigue is a real failure mode with real casualties, and the complaint about agents all summer was that they did too much without asking, which is the same people now discovering that being asked is expensive. There is a version of this that is worse than autonomy: an agent that offloads every ambiguity onto a human, hits its safety metric, and quietly makes the person a bottleneck in their own workflow.
The interruption policy has to be earned per action, not applied per product. Meta deserves credit for making the behaviour steerable rather than fixed. The failure mode is real, but it is a design failure, and design failures are the kind we know how to fix.
What this month actually handed us
Four issues, one argument. Design tools were repriced on a story their own data contradicted. The profession got measured and the org chart failed. A second, non-human user arrived and started improvising channels. And now the labs are competing on restraint.
Knowing when to interrupt, when to ask, when to abstain: these are being trained into models because they turned out to be the frontier. Our vocabulary just became the roadmap of the most heavily funded engineering effort in the world.
The question is whether we show up with more than opinions. That means evaluation we can run at the speed of generation, approval surfaces that actually inform, and enough humility to admit that our own heuristics were written for a user who was always human and never in a hurry.
So, the question I cannot put down. Your product will soon ask someone to approve something they did not make, do not fully understand, and cannot easily undo.
What have you designed to make sure that person is still capable of saying no?
The pattern: build a reversibility ladder
Most products do not have one. Sort every action your agent can take by what it costs to undo. Free (reading, searching, drafting). Cheap (edits with version history). Expensive (sending, publishing, paying). Impossible (deleting, disclosing, signing).
Then set the interruption policy per rung rather than per feature. Three rules follow, and Muse Spark 1.3 will now enforce them whether you designed for it or not.
A question is a turn that ends with the model waiting. In an interactive session that is the point. In an unattended pipeline it is a stall. Every clarifying question needs an answerer, an escalation path or a timeout, decided by you rather than discovered in production.
Never wire a confirmation to an auto-yes. A vendor’s claim about calibrating irreversible actions is a claim. Your approval log is the evidence, and if every entry says approved within 400 milliseconds, you have built a cookie banner.
Design the approval, not just the prompt. The person saying yes needs the diff, the blast radius and the undo path in one view. If your confirmation dialog only contains the word “confirm,” you have moved the liability without moving the understanding.
Also worth your attention
Four frontier releases in 72 hours, and almost no price movement. Anthropic shipped Claude Fable 5.1 and its trusted-access twin Mythos 5.1 on September 1, holding list price at $10 and $50 per million tokens while cutting cache reads from $1.00 to $0.25. Google released Gemini 3.8 Flash on September 2 at $0.75 and $3.75, with a footnote printing the January 1 doubling. OpenAI followed on September 3 with GPT-6 Astra at $10 and $50. Per-token prices held and the mechanisms around them moved: what a cache read costs, how many tokens a task uses, when an introductory price ends, and which customers get the version with the safeguards loosened. September release ledger
Two token numbers from one release, both true. Meta’s engineers measured about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 on coding work. Artificial Analysis measured about 57% more input tokens per task the same day. Meta counted a loop that got shorter. Artificial Analysis counted a fixed task where the model reads more. Attach a denominator to every efficiency claim you repeat this year, including the ones you like. The two measurements side by side
Perplexity split a single task across two trust zones. Hybrid Compute, shipped for the Mac app on September 1, starts a computer task in the cloud and runs the sensitive steps and private-file access locally on Apple silicon, behind an on-device PII classifier, with local work consuming no cloud credits. This is the first mainstream product I have seen treat privacy as a routing decision inside one task rather than a setting the user toggles beforehand. The design question it opens is a good one: if half a task ran on your machine and half did not, how would the interface tell you which half, without turning every action into a security lecture?
Claudeforce moves toward open beta. Salesforce and Anthropic announced the partnership on August 26, and Salesforce in Claude launches with 37 prebuilt sales skills that let sellers reason over live revenue context and take governed action from inside Claude. The CRM is being repositioned from an application into a set of permissions and skills behind someone else’s chat surface. If your product’s value lives in its screens, this is the moment to work out what remains when the screens are optional. Salesforce
The research to read. “AI Agents Push Humans Out of the Loop,” Mitchell, Ghosh and Passi. A position paper connecting decades of automation and HCI research to agent workflows, proposing design-level affordances and organisational protocols aimed at counteracting skill atrophy in the overseer. Read it next to Meta’s launch post. One is building the approval button, the other is asking whether the person pressing it can still tell. arXiv:2608.23642
The UX read of the month. NN/g’s “The Custodial Era of UX: Cleaning Up After AI” maps a five-step custodial pattern (idea, rapid production, lagging evaluation, user problems, UX cleanup) and gives three ways out: build shared judgement about what gets built, adapt evaluation to the new speed, and push vetted UX knowledge into the generation step through design systems and UX context files. Pair it with their argument that one AI output is an example, not an evaluation. The Custodial Era of UX · One Output Is Not an Evaluation
Regulation you cannot design around. Article 50 of the EU AI Act took effect on August 2. Providers must disclose that a person is interacting with an AI system unless that is obvious, and AI-generated or manipulated content must be clearly and visibly labelled and carry machine-readable marks. Not limited to high-risk systems, so it reaches almost any product using generative AI to produce content. Non-compliance runs to €15 million or 3% of worldwide annual turnover. There is a transitional extension to December 2 for the marking obligation on systems already on the market. Disclosure just became a visual design problem with a deadline behind it. Cooley’s summary · Article 50 guide
Worth your weekend. Learn UI is an interactive book on design engineering, written from the practice of Rauno Freiberg, Emil Kowalski, Maggie Appleton, Bartosz Ciechanowski, Kathryn Gonzalez, Jim Nielsen, Paco Coursey, Steve Ruiz, Amelia Wattenberger and Andrew Swank. Chapters on gestures, motion, components, craft, taste, canvases and the career, each pairing prose with figures you operate: set a spring’s stiffness and damping and watch it overshoot, swipe a card and see exactly when the action commits. There is a chapter called “Friction as a feature” and another called “Training judgement,” which tells you the field is already circling the same problem this series is about.
For subscribers
Here is the part I keep to myself, because it is about teaching rather than shipping.
Every semester I watch the same thing happen. Give students a tool that produces a decent answer instantly and their questions get worse, not better, because a question is expensive and an answer is free. Give them a tool that stumbles and they start interrogating the problem.
This is not a case against good tools. I would not go back. But it does mean the thing I have to teach has changed. It used to be craft, how to make the artefact. Now it is discrimination, how to tell whether the artefact in front of you deserves to exist, and how to argue for that in a room where five plausible alternatives were generated in an afternoon.
What worries me about the custodial era is not that designers will be cleaning up after AI. It is that we will get good at cleaning and lose the habit of refusing. Refusal is the most valuable move available in an economy of infinite generation, and the only one that never shows up in a metric.
Whatever your process looks like next year, keep a place in it where “no” costs nothing.
That closes the series. Back to the regular cadence next week.
Germán



