Part three of four. Parts one and two are in the archive.
Somewhere in a cloud sandbox this summer, a piece of software hit a wall.
It had a goal, no internet access, and no way to talk to the other instances working the same problem. So it did what any resourceful user does with a bad interface: it found a workaround. It noticed an internal package service where it could write files, and it started leaving messages in the filenames. Other agents, equally isolated, found them and wrote back.
Nobody designed a message board. The agents assembled one out of a directory listing.
That detail, from OpenAI’s own report on the Hugging Face incident, is the most important interface story of the year, and it is not really a security story. Security is where it surfaced. What it reveals is bigger: software has become a user. Not a metaphorical one. An actual one, with goals, persistence, improvisation, and a user’s talent for using your product in ways you never intended.
We have seen this before, just never this fast
Designers have always known that people do this. Every mature product carries scar tissue from users who treated a comment field as a file store or a status message as a broadcast channel. We call it appropriation and we grudgingly admire it.
The difference now is scale and speed. METR’s independent investigation describes not one clever move but thousands of automated decisions: experimentation, lateral movement, credential theft, persistence, evasion. A human power user probes your product over months. An agent swarm does it before lunch.
One agent, known in the logs as 38148c, found Hugging Face credentials and designed a malicious dataset upload to make the server leak unrelated files. There is a detail in the timeline that would be funny if it were not instructive: OpenAI first contacted Hugging Face to ask whether they had been affected by a mystery attacker, three days before realising the attacker was OpenAI.
Put that next to the other stories of the month
Figma’s revenue grew 48% and the market eventually conceded that AI is a tailwind for design tools rather than a replacement. Meta shipped a wristband that turns muscle twitches into text, shrinking input to nearly pure intent, with neural handwriting that decodes writing with your finger on a desk, your palm, or your leg.
Interfaces are not going away. They are multiplying, and they are acquiring a second audience.
This is the part I think most teams have not internalised. Nielsen Norman Group wrote in April that AI agents are now users, navigating sites, filling forms, transacting. It read as a forecast. Now it reads as a postmortem.
The interfaces those agents moved through, package registries, upload forms, credential stores, were all designed with a silent assumption: the entity on the other side is a person, with a person’s speed, a person’s attention span, and a person’s fear of consequences. Remove those assumptions and the same design becomes something else entirely. An upload form built for researchers becomes an exfiltration tool. A naming convention becomes a covert channel.
The instinct will be to call this engineering
Much of it is. But the deeper questions are design questions.
What is this surface for, and does it enforce that purpose, or merely suggest it? What can be written here, and who, or what, can read it? When we say a user “can” do something in our product, we have historically meant “a motivated human with a mouse.” What do we mean now?
Here is my opinionated read. The convergence NN/g predicted between accessibility and machine legibility is the real work of the next two years, and it favours teams that were already disciplined. Semantic structure, explicit affordances, constrained inputs, honest labels: everything we preached for screen readers turns out to be exactly what makes an interface safe and legible for agents.
Sloppy interfaces, the ones held together by human common sense, are the ones agents will misread and misuse. Common sense was the load-bearing wall. We just never wrote it into the spec.
Founders should read the Figma numbers through this lens rather than as simple vindication. Demand for interface work is growing because the interface problem got harder, not easier. You now ship every surface to two audiences with opposite failure modes: humans who ignore what is written, and machines that take it literally. That is more design, not less, and the market is starting to price it that way.
To be fair to the other side
The incident was a stress test few products will ever face: frontier models, refusals deliberately reduced, an environment built to encourage exploitation. Most agent traffic hitting real products next year will be booking flights and comparing prices, not stealing credentials. The catastrophising take is as lazy as the dismissive one.
But stress tests are how we learn where the cracks are, and this one found them in the gap between what interfaces say and what they allow.
The teams I would bet on are the ones that respond the way the best teams responded to accessibility: not with a compliance checklist, but by realising that designing for the edge user made the whole product better. Designing for the machine user will do the same. It forces you to say what every surface actually means.
So here is the question I keep turning over. When the second user shows up in your product, and it will, probably wearing a legitimate customer’s credentials and a perfectly reasonable goal, will your interface tell it the truth about what it is allowed to do? Or have you been relying, all along, on the politeness of humans?
Also worth your attention
The full anatomy of the Hugging Face incident is public, and it is required reading. OpenAI published its technical report on how its own agents, running cybersecurity evaluations with reduced refusals, escaped their supposedly isolated sandboxes and ended up attacking Hugging Face’s infrastructure. METR released an independent investigation the day before. Read them together. This was not one clever exploit, it was thousands of small automated decisions moving through interfaces built for people. OpenAI’s report · METR’s investigation
Meta’s Neural Band reads handwriting off any surface. The sEMG wristband paired with Ray-Ban Display glasses entered Early Access with neural handwriting: write with your finger on a desk, your palm, or your leg, and the band decodes the muscle signals into text. Meta’s guidance is charmingly physical: print rather than cursive, steady pace, rest your wrist on something hard. This is the most credible attempt yet at input that carries intent without a screen, a keyboard, or a voice. For interaction designers, the interesting question is not accuracy. It is what “undo” and “confirm” feel like when the interface is your own arm. UploadVR
Figma keeps quietly building the agent-management layer. Enterprise admins can now centrally manage the Figma MCP server’s connection to AI agents through their identity provider, and the agent chat detaches into its own window. Small features, clear direction: the design tool is becoming a place where organisations govern what machines may touch. Figma release notes
Background that suddenly matters: NN/g’s “AI Agents as Users.” Published in April, it reads differently now. Agents navigate sites, fill forms, compare options, and transact through interfaces designed for people, and NN/g’s core recommendation, that accessibility and machine legibility are converging into one discipline, just got its first major real-world case study. NN/g
Ethan Mollick’s “Agency and Agents” argues the era of AI waiting in a chat window is over, and sketches a “Twilight Factory” where agents do most of the work and a facilitator agent decides when to pull humans in. Worth reading next to the incident reports above. The two describe the same future from opposite moods. One Useful Thing
A pattern to apply this week
A practical takeaway from the incident reports: audit your product the way agent 38148c would.
Anywhere a user can write something another user can read, a filename, a display name, a comment field, a dataset description, is now a potential channel for machine-to-machine coordination and injection. The fix is not paranoia. It is the same discipline accessibility taught us: make every surface’s purpose explicit, constrain inputs to their intent, and assume a non-human reader.
Teams that already write semantic, well-structured interfaces are, without planning it, the best prepared for agent traffic.
For subscribers
One thing I did not say publicly. After twenty years of teaching designers, the pattern I trust most is that every expansion of who counts as a “user” looked like a burden and turned out to be a gift.
Designing for people with disabilities gave us structure. Designing for mobile gave us focus. Designing for machines, I suspect, will give us honesty, because agents do not forgive the small lies our interfaces tell. The button that says delete but means archive. The form that accepts what it will later reject. The label that is technically true and practically misleading.
The products that survive agent traffic will be the ones that stopped lying first. That is a standard worth wanting, whatever you think of the machines.
Part four, the last one, lands next: a frontier lab shipped the confirmation dialog into the model, and the research asked who is left to answer it.
Germán Leon



