Last issue was about Figma putting code on the canvas, a design company defending itself against AI. This week, the mirror image happened: the most powerful AI company shipped a design problem. GPT-5.6 set records, ChatGPT Work arrived on every plan for free, and the story that dominated the week was not capability. It was an interface that respected nobody’s workflow. This issue is about what that inversion means, because I think it marks the moment the bottleneck officially moved.
The most important story: OpenAI wins the benchmark and loses the launch
On July 9, OpenAI shipped its biggest release since GPT-5: the GPT-5.6 family (Sol, Terra, Luna), ChatGPT Work (an agent that turns a goal into finished sheets, slides, docs, and web apps), and a unified desktop app that merges Codex, Chat, and Work into one surface, available on every plan, including Free.
The capability numbers are real. Simon Willison’s notes report GPT-5.6 Sol scoring 53.6 on Agents’ Last Exam, a new record on long-running professional workflows across 55 fields, beating the nearest competitor by 13 points, with the smaller Terra and Luna models outperforming rivals at a fraction of the cost. Free agentic work for hundreds of millions of users is a structural event.
And yet the loudest story of the week was the app. John Gruber titled his post “Today’s the Day OpenAI F***ed Up the ChatGPT Mac App.” Benedict Evans called the new app “an incredibly confusing sloppy mess” with “the veneer of a polished app without actually being organized or structured or labeled in ways that add clarity and coherence.” TechRadar collected user reactions under the headline “this change could be disastrous.” The specifics are almost comic: the app people used is now called “ChatGPT Classic,” the Codex app became the main app, only the five most recent conversations are visible without extra clicks, and the native Mac app became an Electron package.
Evans’s diagnosis is the part worth saving: at OpenAI, researchers have real influence and product designers do not, and the app shows it. Why it matters: for a decade the model was the bottleneck and the interface was a wrapper. This week the model outran the interface at the exact moment the interface became the product. For practitioners, the implication is blunt. Capability is now rented, interchangeable, and improving on everyone’s behalf. The interface is owned. The layer OpenAI just fumbled is the only layer where a product company cannot be commoditized.
Read: TechRadar on the backlash, Daring Fireball, and Simon Willison’s GPT-5.6 notes.
Product launches and interface innovations
ChatGPT Work, the deliverable machine. Under the messy wrapper, Work itself deserves attention. It gathers context across your apps, breaks a goal into steps, and runs for hours to produce finished documents, spreadsheets, and web apps. For designers, this is the second front in the same war Figma’s code layers opened: interface production without a designer in the loop, now free on every plan. The right response is not anxiety, it is a question: if a founder can prompt a working scenario planner, what is the deliverable only a designer can produce? Read: OpenAI’s announcement.
Figma keeps shipping while the market watches. Code Layers moved into early access this month, and a July 8 update added parallel AI image edits, so teams keep designing while edits run in the background. Small feature, telling pattern: Figma is designing around AI latency instead of making users wait for it, which is exactly the craft OpenAI’s launch lacked. Read: Figma’s release notes.
The next round is already scheduled. Grok 4.5 went public on July 9, and multiple community sources place Gemini 3.5 Pro’s launch on July 17. Treat that date as a rumor, but the cadence as fact: Frontier releases now land weekly. If your product strategy depends on a capability edge, your strategy has a shelf life measured in days.
Research worth your time
“Design Principles for Human-Agent Interaction” synthesizes 106 papers from 2019 to 2026 across CHI, CSCW, IUI, and AAMAS into design levers for agent interfaces: how agent type, interaction stage, and transparency choices shape trust and performance. Useful precisely because this week proved the field’s biggest player is not reading it. Read: arXiv 2606.20630.
“After the Interface: Relocating Human Agency in the Age of Conversational AI” asks what happens to user agency when software stops being navigated and starts being delegated to. Put it next to ChatGPT Work’s launch, and it reads less like a theory and more like a product review written in advance. Read: arXiv 2605.15064.
Regulation: Brussels blinked, but read the fine print
The EU’s Digital Omnibus agreement pushed the AI Act’s high-risk obligations from August 2026 to December 2027 (stand-alone Annex III systems like recruitment and credit scoring) and August 2028 (AI embedded in regulated products). The simplified compliance framework now extends to companies with up to 750 employees and €150 million in revenue. Why it matters for product teams: the deadline pressure eased, but the definition work did not. If your AI feature touches hiring, credit, education, or safety, the classification question (”is this high-risk?”) still determines your roadmap, and teams that treat the delay as permission to stop documenting will pay for it in 2027. Read: Gibson Dunn’s summary.
Funding: a quiet week with a loud niche
Q3 opened slowly (the July 4 week produced about $37 million across three disclosed agent rounds, led by LinqAlpha’s $22 million Series A), but watch the category quietly forming underneath: design tools built for agents rather than people. Glue gives coding agents an infinite canvas to create and iterate on interfaces. The inspector lets designers edit a live front-end like a Figma file and writes changes straight to the codebase. The pattern: the new design tool startups assume the agent is a user too. Read: the Q3 agent funding tracker.
Voices worth hearing this week
Simon Willison, beyond the benchmark notes, criticized the growing fragmentation of ChatGPT into apps and modes and flagged that safety researchers found universal jailbreaks in every round of GPT-5.6 testing. Benedict Evans’s super app critique doubles as an org-chart theory of product quality: show me who has influence, and I will tell you what ships. Both are worth your time this week: Willison and Evans via Daring Fireball.
The essay: Launch week lost to a sidebar
On Thursday morning, July 9, a few hundred million people opened ChatGPT and found that it had been renamed. The app they used every day was now “ChatGPT Classic.” The main app was something else, a merged surface where Chat, Work, and Codex live behind a toggle, where scheduled tasks colonized the sidebar, and where finding a conversation from last week now takes a hunt through “See All.” No migration notice that made sense of it. Just a different piece of software wearing a familiar icon.
The same morning, OpenAI shipped what is by most measures the most capable model publicly available. GPT-5.6 Sol set a record on Agents’ Last Exam, the benchmark closest to real professional work, beating its nearest rival by 13 points. ChatGPT Work, an agent that produces finished documents and working web apps from a goal, is available for free on every plan. On paper, this was one of the strongest product launches in the history of the industry.
The week’s discourse was about the sidebar.
I want to be precise about why this matters, because the easy reading (big company ships bad redesign, users complain, film at eleven) misses the structural point. Redesign backlash is as old as software. What is new is the asymmetry. The model gained more capability in one release than most products gain in a decade, and the interface still managed to be the story. That has never happened at this scale before, and it tells you where the bottleneck now lives.
For years, the deal in AI products was simple: intelligence was scarce, so we forgave the interface. We accepted the chat box, the wall of text, the buried settings, because behind them sat something genuinely rare. Every quarter of model progress made that deal worse for the model makers and better for the rest of us. This week, the deal expired in public. When Sol answers in under a second and builds your spreadsheet while you sleep, the friction you feel is no longer the model thinking. It is the product failing.
Benedict Evans put his finger on the cause, and his sentence deserves to be read as an organizational autopsy: at OpenAI, researchers have real influence and product designers do not. The app is downstream of the org chart. A company whose executives are consumed with the model race shipped exactly what that attention structure predicts: record benchmarks wrapped in an Electron package that renamed its own flagship to “Classic.” Every founder should read that sentence twice, because your product is also downstream of your org chart, and your users can tell who has power in your company by using your app for five minutes.
Here is the founder half of the argument. Capability is now a rental market. You, OpenAI’s competitors, and the fourteen startups in your category all draw from the same pool of frontier models, refreshed weekly (Grok 4.5 last Thursday, Gemini 3.5 Pro rumored for this Friday). Nothing you rent can be a moat. What you own is the surface: the defaults, the information architecture, the recovery paths, the thousand small decisions about what the user sees first and what happens when things go wrong. OpenAI just demonstrated, at maximum public scale, that owning the best model does not protect you from fumbling the layer you actually own. If the best-funded AI company on earth can lose a launch week to interface debt, so can your seed-stage product.
And here is the designer half. This backlash is the strongest evidence of design’s value the profession has received in years, better than any report, because it is a controlled experiment. Hold capability at record levels, degrade the interface, and watch the market’s reaction: the capability disappeared from the conversation. Designers have spent two years being told the model would eat their jobs. What this week showed is narrower and more useful: the model can eat production, the mockups and the screens and the code. It cannot yet eat judgment about how software should feel to a person with somewhere to be. That judgment was the missing ingredient on July 9, at a company that can afford anything.
The research community saw this coming. A paper published on arXiv this spring, “After the Interface,” asks what happens to human agency when we stop navigating software and start delegating to it. The ChatGPT super app is that question shipped as a product: three modes, one toggle, and an agent that works for hours unsupervised. Delegation interfaces need more design, not less, because the user’s understanding of what is happening can no longer be read off the screen. It has to be deliberately constructed. That is a new discipline, and this week made clear nobody owns it yet.
[EXCLUSIVE PARAGRAPH, NEWSLETTER ONLY] When I taught this to my students, I used a rule I called the taxi test. You get into a taxi in a city you do not know. The driver is brilliant, the fastest route every time. But the doors lock themselves, the meter is hidden, and you cannot tell if you are going the right way until you arrive. Nobody calls that a good taxi, whatever the driver’s talent. Most agent products today are that taxi, and the whole discipline of agent UX reduces to unlocking the doors and showing the meter without making the passenger drive. When I score my own AI workflows with ASPIC, the step students skip is always the same one OpenAI skipped this week: they polish what the system can do and never design how a person stays oriented while it does it.
I do not think OpenAI’s app stays bad. They have the money, the talent, and now the motivation; the fixes will come. The lesson is not about them. It is that we have crossed into the period where AI products fail for ordinary product reasons, and succeed for ordinary product reasons, and the extraordinary thing in the middle is table stakes. The bottleneck moved from the lab to the surface. The companies that notice first will hire accordingly.
Which raises the question I keep turning over: if the interface is now the bottleneck, who inside your company has the influence to fix it, and would a stranger be able to tell from your product?
If this issue was useful, share it with a designer or founder who is deciding what to build next. And if someone forwarded it to you: germandleono.substack.com.



