June 23, 2026
Last week, I watched a graduate student use an AI assistant to complete a research exercise. They typed a question. The AI answered. They copied the answer into their document. The process took about forty seconds. It looked exactly like productivity.
It probably wasn’t.
Ethan Mollick, in his recent One Useful Thing post “Choosing to Stay Human,” describes what his research team at Wharton calls “cognitive surrender” — the pattern where people stop thinking about problems and simply let the AI do the work, even when the AI is wrong. In a study of 758 consultants at Boston Consulting Group, half given access to GPT-4, the results were striking: those with AI vastly outperformed those without it. But when asked to solve a problem the AI was known to fail at, the AI users were significantly less likely to get it right. The same elite consultants who had just outperformed everything stopped using their own judgment the moment a confident-looking answer appeared on screen.
“The AI gave them an authoritative-looking answer that happened to be incorrect, and most of them — the same elite consultants who outperformed on everything else — did not catch it.”
Mollick frames this as a challenge of individual intention — a call to be deliberate about which tasks to hand over and which to keep human. That framing is right and important. But it’s incomplete. Because what I keep coming back to is this: someone designed those consultants’ experience with AI. Someone chose the interaction model. Someone decided how the answer would be presented — with authority, without friction, in a way that made the output feel like a conclusion rather than a starting point.
That someone was a design engineer. And that choice had consequences.
The Architecture Problem
A few months ago, Linas Beliūnas published an experiment in his newsletter that points in a completely different direction. Linas built what he calls a startup operating system in Claude — twelve interconnected skills covering everything from idea validation to board management, each encoding real frameworks, real case studies, and real decision trees from founders who had actually done the work.
The result was outputs that were, by his account, “specific, actionable, and often better than what I would have gotten from expensive advisors.” Not because the underlying model was smarter, but because the architecture was deliberately designed to produce depth instead of surface.
“The problem isn’t that AI is stupid. The problem is that surface-level prompting produces surface-level outputs. The solution is architecture.”
This is the sentence I keep returning to. What Linas is describing, though he frames it as a founder’s problem, is actually a design problem. The quality of the output was a function of the architecture surrounding the model, not just the model itself. Each skill wasn’t just a document — it included diagnostic workflows, decision trees, frameworks from practitioners, and anti-patterns that prevented common mistakes. The system interrogated the user before it answered them.
Most AI products are not architected this way. They are designed for frictionlessness, which sounds like a virtue but functions as a trap. An interface that produces an answer in forty seconds is optimized to feel useful. Whether it produces understanding, judgment, or anything that survives contact with a situation the AI hasn’t seen — that’s a different question, and it’s usually not part of the design brief.
What We’re Actually Responsible For
The reason this matters specifically for people working at the intersection of design and engineering is that we are the ones setting these defaults. When we build an AI product, we make hundreds of small architectural choices that collectively determine what happens inside the person using it.
Do outputs arrive with friction or without it? Does the interface invite the user to engage critically, or hand over? Is there any moment where the system says, in effect, wait — do you actually understand this? Or does it just give the answer and let people feel productive?
Mollick notes that the three major AI companies do offer tutoring modes that push users to think rather than receive — but they’re not intuitive to reach. Gemini’s is a tap and a menu. ChatGPT needs you to type /learn. Claude’s is buried in style settings. The default, on every platform, is the answer-delivery mode. That is not an accident. It’s the accumulated outcome of design decisions optimized for engagement metrics and time-to-first-value. The tutoring modes exist; they just lost the internal argument every time.
This is the gap that design engineers should be losing sleep over. Not “how do I make this interface more elegant,” but “what does the architecture of this interaction teach the person using it?” Because the architecture is the answer to that question — not the model, not the prompt, not the user’s individual willpower.
The same research Mollick cites makes the flip side visible. A five-month AI tutoring experiment across ten high schools in Taipei — close to a thousand students — found that a personalized AI that pushed students to solve problems themselves produced gains equivalent to six to nine months of additional schooling, without any added teacher workload. The model was not dramatically better. The architecture was completely different. One system answered. The other one asked.
Designing for the Human Side
The deeper challenge is that we don’t have clear norms yet. As Mollick puts it:“We clearly don’t have rules or role models for how AI gets used at work, in schools, or in government. That’s a problem, but it also means that every organization figuring out a good way to use AI right now is setting a precedent for everyone else.”
The window is open. It won’t stay open indefinitely.
What Linas demonstrated with his startup OS is that the intervention doesn’t have to be subtle or academic. A well-designed architecture can actively produce better thinking — not by lecturing users about their cognitive habits, but by giving the model enough structured context that it asks better questions back. The skills framework he built doesn’t just produce better outputs. It produces outputs that help founders understand why they’re making the decisions they’re making. The system interrogates before it advises.
That’s the design move. Not “give people the answer faster.” Not even “explain the answer while you give it.” Something closer to: design the interaction so that the human who comes out the other side knows something they didn’t know before, and couldn’t just as easily have skipped.
There’s a version of every AI product we’re building where users emerge with sharper judgment. There’s another version where they emerge having outsourced their thinking and now can’t get it back. The difference between those two products is not the model. It’s the architecture. It’s us.
The student I watched last week probably got their homework done on time. The section they filled in was fine. What I still don’t know is whether they learned anything — and whether the person who designed that AI interface ever thought to ask.
That might be the most important design question of the next five years. I’m curious whether the people building these products are treating it that way. Are you?
Sources
Ethan Mollick, “Choosing to Stay Human”, One Useful Thing (May 26, 2026)
Linas Beliūnas, “The One-Person Unicorn”, Linas’s Newsletter (Feb 2, 2026)



