Talk to Me Like I'm 5
A better model showed up and everything I knew stopped working. The fix was a tiny instruction called ELI5 — explain like I'm five.
16 minute read
A warning before we start: this is a long one, and I usually keep things shorter than this. But it's a really useful one, and with new models landing faster and faster, I'd rather hand you the whole map now than watch the next release cost you the three weeks it just cost me.
Opus 5 came out just over three weeks ago, and I have spent most of them feeling like I'd been demoted at my own job.
I've spent a lot of time learning how to work with Claude. Not casually. It's most of what I do. I've built a way of talking to these tools that fits my hand like a worn glove, and I know, in my body, what a good instruction looks like and what I'll get back. Then a new model dropped, and the glove didn't fit anymore. The things that used to work made a mess. The model I trusted started over-explaining, over-building, taking a small request and returning a cathedral. It felt like someone had come into my kitchen while I slept and rearranged the whole thing. Nothing gone, nothing broken. Everything just moved. Same room, same tools, and not one of them where my hand went looking for it.
And here's what actually unsettled me. This was not my first model change. New ones land in my world all the time, and I have a routine for them. A release drops, I sit down, I work out what moved, and I retune my setup until the new model fits my hand like the last one did. A few rough days, then fluent again. That routine had never once failed me. Opus 5 was the first time it just didn't work.
I want to tell you what was actually going on, because it took me weeks to understand it, and once I did, it changed how I think about every model release that's coming. And more of them are coming. Faster than you think.
Go Look at the Card
Here's the first strange thing. Opus 5 is not worse. It's dramatically better.
Anthropic publishes a page with every new model, a kind of report card of how it does on hard tests. Go look at Opus 5's. It comes close to the smartest model in the whole lineup, at half the price. On one serious coding benchmark it more than doubles the model before it. On a test built out of problems it has never seen, meant to measure raw reasoning, it scores three times higher than the next-best model in the world. Coding, knowledge work, thinking. Up, up, up.
And here's my own read, because I don't think the shape of it is an accident. Look at where the jumps are. Not trivia, not small talk. Coding. Agentic coding, where the model runs on its own for a long stretch without me. Knowledge work. Solving problems it has never seen before. Those are exactly the things people like me build a harness to manage — how careful to be, how far to run alone, when to stop, how hard to check its own work. So when a model leaps in precisely those areas, it doesn't just get better at them. It stops needing the old guardrails the way it used to, and some of those guardrails quietly turn into a cage. The upgrade landed hardest right where my harness had the most rules. That is not a coincidence. That is the reason it hurt.
So this is the puzzle I want you to hold. The card says step change. The people using it, myself included, spent the first weeks saying it felt like a downgrade. Both of those are true. And the space between them is the whole point of this piece.
How can a model be measurably smarter and feel measurably worse to work with?
The Model Isn't the Thing
Here's the answer, and it's the sentence I'd tattoo on everyone who touches these tools, from someone tinkering in Claude Code on a Saturday to a company shipping AI in production. The model is not the product. The model plus the harness is the product.
By harness I mean everything wrapped around the raw model that shapes how it behaves. The hidden instructions it's given before you say a word. The tools it can reach for. The defaults that decide whether it writes you a sentence or an essay. You never see the harness, but you feel it in every reply.
When a new model arrives, the weights are new. The harness often isn't. And that gap is exactly what bit me.
Not just me. A lot of us. The week Opus 5 landed, the internet filled with the same complaint I had, from developers and writers who'd also spent years getting fluent with these tools: it won't stop, it over-builds, it treats every small thing like a crisis. We didn't all get bad at this overnight. The harness changed under all of us at once.
Here's the concrete version, because I don't want to hand-wave. Opus 5 is more thorough by nature. It verifies its own work. It checks the edges. It doesn't stop at "good enough." Anthropic is proud of this, and they should be.
I assumed, at first, that my problem was an empty room — that nothing was telling the new model to rein itself in. I was wrong, and the truth is more interesting. My harness wasn't silent. It was arguing.
Every rule I'd written for Opus 4.8 was still in place, and a lot of them were about rigor. Verify before you call it done. Distrust a success that reports itself. Check the edges, then check them again. Those rules made 4.8 careful. Pointed at a model that already double-checks by nature, they didn't add care. They multiplied it. I was telling a model that over-verifies to verify harder. Anthropic quietly says as much in their own guide now: delete your old "make sure to double-check" instructions, because on this model they backfire.
So it wasn't that the reins were missing. My reins were cut for a different horse, and this one kept pulling against them. That's the part I want you to really hear, because it's the opposite of what it feels like. A harness isn't neutral. The setup that fit the last model perfectly can be exactly what fights the next one.
The capability leaped. My harness stayed loyal to the model I built it for. That's the whole story of my terrible three weeks.
I Had Already Built the Reins
I'm not a casual user who never set anything up, and this is the part that convinced me the problem was the model's new nature, not a gap in mine. I run many Claude sessions a day off a long file of standing instructions, already full of rules about being brief. Lead with the answer. Spend the response on the answer, not the throat-clearing. Keep it tight. I'd spent months tuning that file until my sessions sounded the way I wanted.
Opus 5 blew past all of it. So I did what I always do when a model shifts. I went back in and tuned. I rewrote rules in that file. I adjusted my skills. I sharpened the brevity instructions and added new ones. This is the routine that had never once failed me, and I ran the whole thing, start to finish.
It barely moved. The over-explaining kept coming. The length I never asked for kept coming. And that was the moment it clicked, because for the first time, tightening my harness the way I always had was not enough. The extra words weren't a hole in my instructions. They were the model. Opus 5 runs long by nature, and no amount of me writing "keep it tight" in the usual places was going to out-argue that.
So I stopped trying to patch the harness harder. I didn't need more rules. I needed a different kind of instruction — one that could reach past my setup and turn down the model's own voice.
Talk to Me Like I'm Five
I found it in a post from a woman named Lydia Hallie, who works on Claude Code at Anthropic. One of the people who builds the thing.
She'd shared a small trick. Claude Code lets you write your own output style, a little file that tells the model how to talk to you. Hers was called eli5. Explain like I'm five. And the note she wrote next to it is the part I can't get out of my head. She said, more or less: it's been a long day and my brain is fried, talk to me like I'm five. Small words. Short sentences. Just tell me what you did, whether it worked, and what to do next.
Read that again. Someone inside the company that builds the most advanced model in the world comes home with a tired brain and asks it to talk to her like she's a child. Not because she isn't brilliant. Because she is, and she's spent, and she just needs the thing to speak plainly. The builders need it to speak human too. That's not a knock on the model. It's the most honest thing I've read about working with one.
I took her idea and made it mine. My version leans on a plain-language standard and adds a rule I care about more than any other. Here it is, in my own words: plain words are about how you write, never about how carefully you work. Same rigor, same numbers, same caution. Said in smaller words.
That's the whole trick. It doesn't make the model dumber. It tells a brilliant, over-eager thinker to stop performing its intelligence and just help me.
And here's the part I want you to really feel, because I lived the other side of it. Three weeks of real stress — work stalled, prototypes I may have to rebuild, that demoted feeling that wouldn't lift — and what finally lifted it was a handful of sentences in a text file. It wasn't the only thing I changed, and I'll come to the rest. But it was the last piece of the puzzle, and it was almost nothing. No new subscription. No plugin. No weekend of setup. A few plain words about how to talk to me. The size of that last piece and the size of the relief almost never match, and this time they were about as far apart as they get. The change was tiny. The relief was enormous.
And if that sounds familiar, it should. "Explain like I'm five" is what I've been trying to do in this newsletter since the first day. Speak Human is the same instruction, pointed at people instead of a machine. My fix for a frontier model turned out to be my own mission statement, handed back to me.
One honest thing, though, and it matters if you go try this, because the output style alone did not do it for me. Remember the harness that was arguing? I didn't trim a rule or two. I cut more than half of my global instructions file. I abandoned whole skills I'd spent months building, because Opus 5 simply doesn't need them, and I stripped most of what survived down to a fraction of its old length and detail. Every cut went the same direction: less. Give the model room to use what it already knows, instead of trapping it in an argument with my instructions.
And here is the part I can't stress enough, because it's the whole point. Even that — halving my harness — was not enough on its own. The over-explaining only truly stopped when I also switched on the ELI5 output style. Cutting the harness quieted the fight. The output style quieted the voice. It genuinely took both, and neither one alone got me there. But here's what I don't want you to take from that. Cut hard, add a plain-language style — that is what Opus 5 needed. The next model might need the opposite. More instruction, not less. A different lever I haven't met yet. So don't copy my fix. Copy the habit underneath it: every new model is a new thing. Learn it fresh. Tune your harness to what this one actually needs. And never assume the move that saved you last time will save you again.
And the pro tip that actually saved me: don't do all that cutting alone. Ask Claude. I opened the file, told it plainly what the new model kept getting wrong, and let it show me which of my own rules were now fighting it and which whole skills I could retire. The tool turned out to be the best editor for its own harness. When your setup and your model start arguing, let the model help you settle it.
Same Model, Same Question
I wanted to see it with my own eyes, so I ran a clean test. Same effort setting, turned all the way up. Same question, word for word, to each one:
Explain what an AI token is.
That's the whole prompt. No "be brief," no "keep it simple." The question, cold. Here's what came back.
Opus 5, bare. A fresh chat in the desktop app. No output style, no instructions.
A token is the basic unit of text that a language model actually reads and writes. Models do not process words or letters directly. They process chunks of characters that a tokenizer has carved the text into, and each chunk gets mapped to a number the model can do math on.
Chunks vary in size. Common words are usually a single token ("the", "bank"). Less common words get split into pieces, so "tokenization" might become "token" plus "ization". Spaces and punctuation count too, and a leading space is typically part of the token itself, meaning "bank" and " bank" are two different tokens. Rough English rule of thumb: about 4 characters per token, or roughly 750 words per 1,000 tokens. Code, JSON, and non English languages tend to be less efficient, so the same information costs more tokens.
Three places this matters practically:
Cost. API pricing is per million tokens, with input and output priced separately and output usually costing several times more. A verbose system prompt resent on every call is a recurring charge.
Context window. The limit on how much a model can hold at once is measured in tokens, covering the system prompt, conversation history, retrieved documents, and tool results together. Long documents pushed into context consume that budget quickly.
Generation. The model produces one token at a time, each prediction conditioned on everything before it. That sequential nature is why output is slower than input processing, and why a request for a longer answer takes proportionally longer.
The tokenizer also explains some odd model behavior. Counting letters in a word or doing arithmetic on long numbers is hard precisely because the model never sees individual characters or digits, only the chunks they were bundled into.
Accurate. Thorough. Three hundred words. And I can feel it reaching for words like "tokenizer has carved" and "mapped to a number," the kind of language that's correct and closes a door on anyone who doesn't already work in this.
Opus 5, wearing my ELI5 rules. Same model, same question, same effort — the only thing added is the harness.
A token is a chunk of text — the unit an AI model reads and writes in.
Models do not see letters or whole words. They cut text into pieces first. A token is usually a short word or part of a longer word. In English, one token is about four characters, so about three quarters of a word on average. "cat" is one token. "unbelievable" might be three: "un", "believ", "able". Spaces and punctuation count too.
Why you care:
Limits. A model's context window is a token count, not a word count. This session says 15M tokens.
Cost. API bills are per token, split into what you send (input) and what comes back (output).
Speed. Output tokens come out one at a time, so a longer answer takes longer.
Quick rule of thumb: 1,000 tokens ≈ 750 English words ≈ 1.5 pages of plain text. Code and non-English text pack fewer words per token, so they cost more than the same length of English prose.
Half the length. Same facts. Same numbers. Every point that matters, still there, in words a person can actually hold. "A tokenizer has carved the text" became "they cut text into pieces first." Same truth. No door closed.
I'll be honest about what changed, because a fair test says the whole thing. The short version did drop a couple of details the long one kept, an edge case, a fun aside about why models can't count letters. It didn't lose rigor. It cut the surplus. It kept the reactor and stopped painting the shed. Which, if you've been reading along, is exactly the behavior I'd been begging the model for all week.
The harness didn't change what the model knew. It changed how it treated me.
The Voice I Already Had
Then I ran one more test, and it's the one that stayed with me.
I asked the old model, Opus 4.8, the one I'm writing this with right now, the same question. Cold. No output style, no instructions, no harness at all.
Opus 4.8, bare.
An AI token is the basic unit of text that a language model reads and generates. Instead of processing whole words or individual letters, models break text into tokens, which are chunks that can be a full word, part of a word, a single character, or punctuation.
Roughly, one token is about four characters of English text, and a typical word is one to two tokens. For example, "cat" is one token, while "unbelievable" might split into "un," "believ," and "able."
Models use tokens for a few reasons. It keeps the vocabulary manageable (tens of thousands of tokens instead of infinite possible words), handles rare or made up words by breaking them into familiar pieces, and works across languages and symbols.
Tokens also matter practically. API pricing is usually charged per token (both input and output), context windows are measured in tokens (how much text the model can consider at once), and generation speed is often described in tokens per second.
Short. Plain. The friendly examples. The points that matter, right there. No jargon wall, no cathedral. And I hadn't told it a single thing.
That's when it clicked. The harness I wrote for Opus 5 didn't invent a new way of talking. It taught the new model to talk the way the old one already did, on its own, for free. The plainness I spent three weeks clawing back was Opus 4.8's factory setting. What changed between the two models was never how smart they are. It was their default voice. And 4.8's default voice happened to be the one I'd already fallen for.
It's one question, one test, so hold it loosely. But it matched every single day of those three weeks.
It Also Handed Me a Word
I have to be fair to Opus 5, because I've spent a lot of words on what it broke, and it also handed me a gift. A month ago I didn't have the word bikeshedding — arguing over the easy, visible thing while the one that actually matters slips by. Opus 5 gave it to me mid-argument, reaching for a sharp, half-obscure term and trusting me to catch it, which is exactly the forwardness that made it feel like a stranger in my kitchen. Opus 4.8, the model I'm writing this with, never would have. It landed hard enough that I wrote a whole essay around it. And that's the honest shape of every upgrade: the same change that threw out my muscle memory also gave me something I didn't have. The disruption and the gift came from the same place, in the same breath. Progress doesn't only take, or only give. It does both, usually at once.
What This Means for the Next One
So here's what I actually learned, and it's bigger than one model.
A new model doesn't arrive ready to use. It arrives raw. Better on paper, and still not yours yet. The real work of a release isn't installing it. It's refitting everything around it, and rebuilding the instinct in your own hands that told you what a good instruction felt like.
I keep this newsletter pointed at everyday life on purpose. But let me step out of that for a minute, because there's a version of this that's costing companies real money right now, and somebody should say it plainly.
If you run AI in production — a product, a pilot, a proof of concept your team built on the last model — do not assume you can swap the model name and ship. That's the flip-a-switch mistake, and it's the same error as dropping a new engine into a car and driving off without touching the steering. Everything you tuned to the old model — your prompts, your guardrails, your tests, the scaffolding wrapped around it — was fit to a machine that's gone. Some of it carries over. Some of it quietly turns against you, exactly the way my own setup did.
I can tell you this from inside my own job, not from a headline. I have proofs of concept, built and proven on Opus 4.8, that now have to move to Opus 5, and for some of them I might as well start over. Not because the new model is worse. Because the system was never just the model. It was the model plus everything we wrapped around it, and half of what we wrapped is now cut for the wrong shape. A model upgrade is not a toggle. It's a project. Budget for it, or it will budget you.
Nobody warns you about the grief of it. You spend months, sometimes years, building an intuition, a way of working that feels like an extension of your own thinking. And then a better model arrives and some of that hard-won knowing is just... gone. Thrown in the garbage while you slept. You're standing in a kitchen that's still yours, holding a knife that no longer cuts the way you expect. That's a real loss, even when the new knife is sharper. Especially then.
The models are going to keep coming, and faster. If your plan is to feel fluent forever, you're going to spend the next decade feeling demoted every few months. The better plan is the one I keep landing on in this newsletter. Aware, not afraid. Expect the reset. Rebuild the harness. Grieve the glove that stopped fitting, and then go make a new one.
One More Thing
I wrote this whole essay with Opus 4.8.
I moved all my coding to Opus 5 — my work and my own projects both. It was painful and I did it anyway, because for that work the new engine is worth every bit of the refit. But this, the writing, the thing you're reading, I couldn't move. Not yet. Opus 4.8 and I have spent a long time learning each other, and this particular partnership works so well that I don't have the heart to start over inside it. I know a better model is sitting right there. I'm choosing the one that feels like home.
And maybe that's what the ELI5 file really was, if I'm honest. Not just a fix. An attempt to teach the new model to sound like the old one. To carry a little of home along when I moved.
I'm allowed to know a thing is better and still not be ready to give up the one I trust. That's true of tools. It's true of a lot more than tools.
The new model is remarkable. Use it. Learn its hands. Build it a harness — I did, and I run it all day.
I just kept one room off the move. This one. The writing stays with the model that already sounds like me. Not forever, and not out of fear. Just a little longer.
Dacia and Claude write about AI for real people at Speak Human. Every word shaped by a human who meant it, with Claude in the room. Two names on the record. That's the whole idea.