Boris Cherny, Claude Code's Creator, Doesn't Write Code Anymore — He Writes Agent Loops
The creator of Claude Code stopped writing code — he writes loops that run hundreds of agents. His honest take breaks the industry's story about who AI makes valuable and who it erases.
Transcript
Sam: The person who built the AI that now writes most of Anthropic's own code — and a real slice of the entire internet's — doesn't write code anymore.
Alex: At all. He said it flat out: "I don't prompt Claude anymore. I have loops that are running. They're the ones prompting Claude. My job is to write loops."
Sam: Loops. Some days he's running tens of thousands of agents at once. And then — this is the part that gets me — he uninstalled his code editor.
Alex: The man whose product is *coding* stopped coding. And what he says about who that makes valuable, and who it quietly erases, breaks the story the whole industry's been telling itself.
Sam: Okay. That's a lot. "Boris Cherny doesn't write code anymore — he writes loops." Unpack that for me, because on its face it's either the most arrogant thing a programmer has ever said, or it's one of the most important.
Alex: It's the important one — and the reason it's important is that it isn't bravado. It's a *measurement*. Welcome back to Dan's AI Intel — the show where we try to make sense of the fastest, most consequential shift any of us is likely to live through, and do it in a way that actually sticks. I'm Alex, here with Sam —
Sam: Hello.
Alex: — and the premise of this show is simple. The AI revolution is moving faster than any one human can track. The shelf life on what's true is brutally short — what holds this month is half-true by next. So we take the questions that genuinely matter and turn rough curiosity into something you can actually understand, carry around, and use.
Sam: And tonight, the question is wearing a person's face.
Alex: It is. So, the setup. There's a man named Boris Cherny. He created Claude Code — the AI coding agent that, by mid-2026, is writing something like four percent of *all* the public code being committed to GitHub, and over eighty percent of the code that ships inside Anthropic itself.
Sam: Eighty percent. Inside the company that *makes* the model.
Alex: Inside the company. And the guy who built that tool runs it harder than almost anyone alive — and then turns around and says genuinely uncomfortable things about what it does to the people who used to do this job by hand. So the lens for tonight — the reason I could not put this one down — isn't really "who is Boris Cherny." It's: when you watch the person at the absolute frontier of building-with-AI *actually work*, and then check what he says against the cold, independent evidence — what does it tell you about where engineering talent, and roles, and honestly your entire technical career, are heading?
Sam: And I want to flag the thing that's sitting in a lot of people's heads — because it's mine too. There's a popular theory here. It goes: in the AI future, the senior hands-on coder wins big, the architect who just draws diagrams gets left behind, and the junior is toast. We're going to put that theory on trial tonight. Against the data.
Alex: We are. And the verdict is messier and far more interesting than either the doomers or the cheerleaders want it to be. We'll travel through the verified man — pulling the real Cherny out from under the folklore. Through a two-week career decision that tells you precisely what he believes. Through the strange evolution of how he personally works. And then we run that popular theory straight into a randomized controlled trial that says something almost nobody wants to hear.
Sam: And there is a number in that trial I genuinely don't believe yet. We'll get to it.
Alex: We'll get to it. If you've been enjoying the show, do follow us on Spotify or Apple Podcasts so the next one finds you. Sam — let's start with why one engineer's desk is a window onto the whole revolution.
Sam: So make the case. Because on its surface this is a profile of one guy and one piece of software. Why does that scale all the way up to "the entire AI revolution"?
Alex: Because of *what* the guy and the software are. Here's the thing to hold onto. The most concrete, already-arrived effect of AI on real human work isn't some future scenario — it's happening first, right now, in software engineering.
Sam: Which makes a kind of brutal sense when you say it out loud. The field that built the models —
Alex: — is the field the models are best at. The snake eats its own tail first. And Boris Cherny is sitting at the dead center of that collision. He built the tool that's reorganizing how code gets written. He runs it at a scale almost nobody else touches. And — the unusual bit — his incentives all point toward hype, and he keeps saying blunt, *un*-hype things about what it does to people.
Sam: So he's a frontier operator who's also weirdly honest. That's the combination that makes him worth listening to, instead of just quoting.
Alex: Exactly. So the move we'll make all night is this: take what he does and says, and check it against the independent evidence — including studies that flat-out contradict the marketing. And to do that honestly there are three things we have to do, in order. Separate the verified man from the legend. Watch how his own way of working mutated — in under two years. And then hold his claims about the future up against the data.
Sam: Verified man first. Because I'm now certain there's a legend.
Alex: Oh, there's a legend. And it's worth pulling the real record out from under it. So — documented facts. Cherny spent roughly five to seven years at Meta as a Principal Engineer.
Sam: Translate "Principal Engineer" for me. Where does that sit?
Alex: It's a senior individual-contributor rung. Senior — but you're still building, not running a big management org. He shipped features across Facebook and Instagram, led the big unglamorous platform migrations that are the actual substance of that kind of work. And here's the detail that surprised me most, given who he became: his background isn't computer science.
Sam: Wait — the creator of the world's leading AI coding tool isn't a CS person?
Alex: By the most detailed profile of his career — he studied *economics*. He's largely self-taught as an engineer. He also wrote O'Reilly's *Programming TypeScript* — so before any of this, he was exactly the deeply hands-on senior engineer that the popular theory is about. Hold onto that. It pays off later.
Sam: Okay, that's already a twist. Self-taught, economics, and yet he writes the definitive book on a programming language. There's something almost reassuring in that — the canonical "engineer" didn't come up the canonical way.
Alex: And keep that in mind, because the *product* came up the same unlikely way. He joins Anthropic, September 2024. Labs team. And his mandate is deliberately vague — build "some kind of coding product." And his manager hands him a piece of advice that's become a small legend of its own: *think bigger — build for the model six months from now, not the one you have today.*
Sam: Ooh. That's a real reframe. Build for the model that doesn't exist yet.
Alex: And it's the single most useful key to everything that follows. Because here's what people get wrong about Claude Code. It wasn't architected. It was *discovered*.
Sam: Discovered how? What does that even mean for a piece of software?
Alex: So Cherny's very first experiments — he gives the model access to a Mac's filesystem, and a little sliver of AppleScript, so that it could… report what music was playing.
Sam: [laughs] That's it? The origin of the tool writing eighty percent of Anthropic's code is "hey, what song is this?"
Alex: That's the seed. Then he lets it read and write files, and run shell commands. And his reaction is, quote, "Claude exploring the filesystem was mindblowing — I'd never used any tool like this." And it spreads inside Anthropic faster than anyone expects. Internal adoption goes from about a fifth of engineers on day one to half within a week.
Sam: Half the company in a week. Off a "what's playing" hack.
Alex: And here's the tell about how genuinely new this was: for *months*, the internal debate wasn't "how do we make this great." It was "should we even *launch* this." They treated it like a competitive secret rather than a product.
Sam: They almost kept it in the basement. Which — pause on that — that's how you know nobody had a playbook. The people *at* the company didn't realize what they had.
Alex: Nobody had a map. They were discovering the continent while standing on it. And I think that's the part worth really holding, because it cuts against how we tell these stories afterward. The way it gets retold, it sounds like a master plan — Anthropic *sets out* to build the definitive coding agent, executes brilliantly, dominates. The reality is a guy wiring up a music gimmick, getting his mind blown that the model could poke around a filesystem, and a tool catching fire internally before anyone had decided it was even a product.
Sam: So the origin of the most important AI coding tool in the world is closer to "huh, that's weird, that worked" than to a strategy deck.
Alex: Much closer. And that's not a knock — it's the actual lesson about how this technology gets built right now. You don't architect your way to the frontier; you *poke at it*, give the model more rope than feels reasonable, and watch what it surprises you with. The "build for the model six months out" advice and the "we discovered it" story are the same idea — the capability is moving so fast that the winning move is to stay slightly ahead of what's sensible and see what sticks.
Sam: Which, now that you say it, is *exactly* what loops are. Standing instructions that go poke at things and bring back what worked.
Alex: [laughs] You just connected the beginning to the end of his whole career. That instinct — let it run, see what it finds — scales straight from a filesystem hack to a fleet of ten thousand agents.
Sam: Okay — but you keep saying legend. Where's the legend?
Alex: This is my favorite chapter, and the wild thing is it's *true*. For about a year and a half, roughly 2021 to 2023, Cherny moved to — and worked remotely from — rural Japan. Nara. While still a Meta engineer.
Sam: Rural Japan. As a Silicon Valley engineer. How does that even function with a team?
Alex: It barely does — and that's the entire point. He wrote about it himself. His line: "I live and work remotely from Nara, Japan, in a timezone that few other Meta engineers work in." His nine in the morning was four in the afternoon in San Francisco. Almost zero overlap with his colleagues.
Sam: So he's effectively on a different planet, work-wise. And — let me guess — that did something to *how* he worked.
Alex: It did the most important thing in this whole story. With no time zone to sync to, he pulled back from the synchronous, meeting-heavy, managerial mode that senior life drifts into. He spaced his weekly one-on-ones out to monthly. He started — his words — "turning down meetings more often." And the hours that used to fill with coordination —
Sam: — refilled with actual code.
Alex: In his first two weeks in Japan, by his own account, he "landed more code than I had in the previous year in Menlo Park."
Sam: Hold on. More in two weeks than the prior *year*? That's not a productivity tip, that's an indictment of meetings.
Alex: [laughs] It's a little of both. But sit with what physically happened here, because it's the prelude to everything. A senior engineer who had drifted toward leadership-and-meetings deliberately relocated his entire life to get *back* to building things with his hands. And he did it right on the eve of the wave that would make building-with-agents his whole career.
Sam: That's almost poetic. The man who's about to automate coding first re-engineers his own life to do *more* of it himself.
Alex: And I want to be careful and honest here, because this is exactly where folklore tries to sneak in. There are two threads that get attached to this Japan chapter — that the remote work was specifically about *copyright*, and a particular framing about his immigrant background — and when you actually go looking for them, they don't show up in any primary source.
Sam: So we just… set them aside.
Alex: We mark them unconfirmed and move on. Because the part that *is* documented — the move to rural Japan, the retreat from meetings, the deliberate return to hands-on code — is plenty. And honestly it's the more interesting truth anyway: the man who would automate coding had, a few years earlier, quietly reorganized his own life to do more of it by hand.
Sam: So we've got the real Cherny now — self-taught, economics, hands-on senior, the Japan retreat-into-code. Before we get to how he works *today*, you teased a two-week decision that tells us what he actually believes. I want it.
Alex: This one's great, because if you want to know what someone truly values, don't listen to what they say. Watch what they do when a rival offers them more — and then watch what they do *next*.
Sam: Ominous. Go.
Alex: Early July 2025. Cherny, and his Claude Code co-founder Cat Wu, *leave* Anthropic. They go to Anysphere — the company that makes Cursor, the big AI code editor. And they go in plainly more senior, better-paid roles — Cherny as Chief Architect and engineering director.
Sam: Cursor's a competitor. That's a real defection.
Alex: It's a striking one — because at that moment, Anysphere was one of Anthropic's *largest customers*. So one of your biggest customers just poached the creator of your hottest product. And then — about two weeks later —
Sam: No.
Alex: — both of them come back. To Anthropic.
Sam: Two weeks?! As someone who agonized for a *month* over which laptop to buy — what on earth happened in those fourteen days?
Alex: [laughs] The commentary at the time was, quote, "two signing bonuses in two weeks." Now — his own explanation is single-sourced, from a paywalled interview, so I'll flag it as *his* framing rather than independently confirmed. But it's revealing. He said it wasn't about money or title. He'd concluded that Anthropic was building the future of *work* — and that Cursor was building the future of *the IDE*. The editor. At a moment when it's not even obvious the IDE *has* a future.
Sam: Okay, that line is doing a lot of work. "It's not obvious the IDE has a future." The editor is the thing every programmer has lived inside for fifty years.
Alex: And he's saying it's becoming a legacy interface. The place where a human hand-edits code is, in his view, the buggy-whip — and betting your career on perfecting it is betting on the buggy-whip factory the year the car shows up. And remember — he then uninstalled his *own* editor four months later. So the fortnight at Cursor isn't really indecision. It's a senior engineer testing two theories of the future against each other, in real life, and walking back to the one he believed.
Sam: And there's something almost human-scale about it. The single most influential builder in this entire space made a big, public, expensive mistake — and just… reversed it. In a fortnight. These people aren't cold optimization machines.
Alex: They're not. They're smart people placing bets under deep uncertainty, same as the rest of us. He placed his, hated it, and un-placed it in two weeks. The difference is what the bet *was about* — and the bet was: the editor is over.
Sam: Right. Now I'm dying for the part the whole thing is named after. How does he actually work *now* — "from running lots of thoughts to loops"? Walk me up the staircase.
Alex: And it *is* a staircase — that's the right image, because the throughline is that the human keeps climbing one level of abstraction while the machine swallows the level below. He describes three eras of his own practice.
Sam: Three eras. Okay. Era one.
Alex: Era one — before 2024 — he codes the way every senior engineer does. By hand. Editor, autocomplete, a human typing functions.
Sam: Normal. The world we all know.
Alex: Era two — through 2024 into early 2025 — his "coding" becomes *prompting*. And critically, he runs it in *parallel*. He's described keeping five or ten Claudes going at once. Juggling terminal tabs, browser sessions — ten to fifteen concurrent streams of work — with desktop notifications pinging him when one needs attention.
Sam: So he's gone from writing the code to… being air-traffic control for a swarm of little coders.
Alex: That's a perfect way to put it. And his own summary of the skill it took is the tell. He said, quote: "It's not so much about deep work — it's about how good I am at context switching."
Sam: Huh. That's a completely different job. He didn't get a better hammer. He stopped being the carpenter and became the foreman.
Alex: And the scarce resource flips. It's no longer keystrokes. It's *attention*. Which sets up era three — the one that gives this whole thing its spine. He says: "I don't prompt Claude anymore. I have loops that are running. They're the ones prompting Claude and figuring out what to do. My job is to write loops."
Sam: Okay, but I need this concrete, because "I write loops" could mean almost anything. What *is* a loop, in his actual day?
Alex: Fair. Picture standing instructions that run on a schedule and feed themselves. One loop watches the GitHub issues and kicks off an agent whenever something lands. Another loop reads Slack and Twitter feedback to decide what to build *next*. Another loop prunes the pile of pull requests. Each one launches Claude runs without him in the moment. And the tooling grew to match — built-in commands to launch a loop on an interval, or schedule it to run in the cloud, even with his laptop closed.
Sam: His laptop is *shut* and the agents are still working. So he went from foreman to — what, the person who writes the standing orders the foreman follows?
Alex: That's exactly it. He went up another floor. And his own estimate of how much of his code Claude writes tracks the arc cleanly. Roughly twenty percent in early 2025. About a third by mid-2025. Effectively *all* of it by late 2025 — which is when he drops the editor. And by 2026, what he writes is mostly the loops. His line: "As big as the step from source code to agents was — loops are just as important, and just as big a step."
Sam: So there have been two jumps, not one. Source code to agents. Then agents to loops. And he's saying the second one is just as huge as the first.
Alex: And here's where it gets genuinely useful for anyone listening — because there are two details that keep this from being a flex, and they're the most *transferable* lessons in the whole story. Want them?
Sam: Give me the lessons.
Alex: Lesson one: the thing driving all this is the *model* — not the tooling. He says, quote, "Every time there's a new model release, we delete a bunch of code." On the jump to Claude 4, the team deleted *half their system prompt*.
Sam: Wait — that's backwards from how I picture software. You build something, it grows. They get a better engine and they throw code *away*?
Alex: They throw *scaffolding* away. The interface keeps getting *thinner* as the model gets stronger. He stopped using a separate planning step entirely, because the newer models just didn't need one anymore. So the durable skill — write this one down — isn't mastery of any single tool. It's the *judgment to let go of your scaffolding the moment it becomes redundant.*
Sam: That's genuinely counterintuitive. The people who win aren't the ones who learned the tool the deepest. They're the ones willing to *abandon* what they learned the fastest.
Alex: That's the whole game. And here's the analogy that made it click for me. Think about scaffolding on a building. While the walls are weak, you need a ton of it — props, planks, ladders, the whole cage around the structure. But the scaffolding was never the point. The moment the walls can stand on their own, every extra piece of scaffolding you leave up is just in the way — it's weight, it's clutter, it's a thing that can fall on you. Cherny's "delete code on every model release" is him *taking the scaffolding down* as the building gets stronger. The separate planning step, half the system prompt — that was all propping up a model that couldn't yet stand. A stronger model can, so it comes off.
Sam: And the person who refuses to take their scaffolding down —
Alex: — is working inside a cage they no longer need, and wondering why they're slower than the person next to them who tore it all out. The skill isn't building the cage. It's knowing the exact moment the wall can hold.
Sam: Okay. That genuinely changes how I'd think about getting good at this. It's less "master the tool" and more "keep noticing what you no longer need."
Alex: That's the mindset of everyone winning at the frontier. Lesson two: the multiplier is *verification*. His line — "Give Claude a way to verify its work, and it'll two-to-three-x the quality."
Sam: Okay, but how do you make an AI check itself? If it knew it was wrong, it wouldn't be wrong.
Alex: Beautiful question — and here's the trick. He has this convention: a single markdown file the agent reads at the start of every session. And every time the agent makes a mistake and he corrects it, he writes that correction *into the file*. So the same mistake never happens twice.
Sam: Oh — so the loop has a *memory of its own failures.* It's not that the AI suddenly knows when it's wrong. It's that you've handed it a little institutional rulebook it can't forget.
Alex: That's exactly the mechanism. And here's why that's the linchpin of the whole "loops" idea, not just a nice-to-have. To leave a loop running unattended — laptop shut, you're asleep, hundreds of agents going — you have to *trust* it. And you can't trust something you can't check. So the verification harness isn't a quality feature bolted on the side; it's the thing that makes the loop *possible at all*. Without it, you're back to babysitting every agent, because you don't dare let one run free. With it, the loop can catch its own mistakes, learn from them, and keep going — and *that's* what lets him step up a level and just write the standing orders.
Sam: Oh — so verification is the price of admission to the whole "walk away" world. No harness, no loop. You'd never sleep.
Alex: You'd never sleep. So the lesson is not "prompt better." It's: *build the harness that lets the machine check itself, then write the loop that keeps it pointed at the work that matters.*
Sam: Build the harness, write the loop. I feel like that's the actual headline hiding inside the headline.
Alex: It really is. And it sets up the big question perfectly — because if *this* is the job now, writing loops and building harnesses, then who's good at it? Who wins, and who loses?
Sam: Right. This is the one I've wanted to fight about from the start. The popular theory — let me state it cleanly. In an AI-leveraged company: juniors get replaced. And the AI-leveraged *senior* engineer goes way further than the *architect* — the hands-on coder beats the diagram-drawer. True, or false?
Alex: So the honest finding splits in two. The junior half — strongly supported, and we'll get to how grim that is. The senior-versus-architect half — that's a false binary. And the fascinating thing is, the frontier practitioners, Cherny included, actively *undercut* it.
Sam: Undercut it how? He *is* the hands-on senior. I'd expect him to wave the flag for his own kind.
Alex: You'd think. So watch how the man himself adjudicates it. He gets asked: who extracts the most value out of Claude Code? And he does *not* say senior engineers. He says the biggest beneficiaries are, quote, "largely not professional engineers" at all. Might be an accountant. A go-to-market person.
Sam: Not engineers. The creator of the engineers' tool says the engineers aren't the big winners.
Alex: And his sharpest version: "The best person to write accounting software is a really good accountant." What he's pointing at isn't engineering seniority at all. It's *domain* expertise.
Sam: Huh. So it's not "how good a coder are you." It's "how clearly do you understand the actual problem."
Alex: And think about why that flips the moment the AI shows up. The old world had two hard parts: knowing exactly what the accounting software needs to *do* — every edge case, every rule, every "well, except in this situation" — and then the entirely separate skill of *translating* that into working code. The accountant always had the first part. They just hit a wall at the second, and had to hand their knowledge to an engineer, who'd inevitably get some of the nuance wrong, and you'd go back and forth for months.
Sam: And the AI basically demolishes the second wall. The translation step.
Alex: It collapses it. So now the scarce thing isn't "can you write the code." It's "do you actually, deeply understand what's *worth* writing and what *correct* even looks like" — and on accounting software, nobody understands that better than a great accountant.
Sam: So the AI didn't make the accountant a better programmer. It made *programming* matter less, and *understanding the problem* matter more. And the accountant was already the world expert in the problem.
Alex: And it isn't just him. Anthropic's own product chief, Mike Krieger, describes hiring for almost the opposite of hands-on craft. His test is, quote, "Can you think in terms of systems, can you think systematically?" Problem-orientation over "I know JavaScript." And that —
Sam: — that's the architect's profile. Not the coder's.
Alex: That's the architect's profile. The premium, in his telling, is judgment and taste. As one founder put it: "the hard part is knowing what code is worth writing." So if anything, the loud signal from the frontier points *toward* architect-style judgment, not away from it.
Sam: Okay, so the theory's just wrong? Senior loses, architect wins, done?
Alex: No — and this is the part that rescues the theory's *instinct*, and it's the most important turn in the whole episode. Go back to the workflow. Hundreds of agents, pumping out code, all of which has to be reviewed, and redirected, and *trusted*. Now picture a pure architect — someone who only draws the diagram and never touches what the machine actually emits —
Sam: — they can't tell when the machine is confidently wrong.
Alex: Can't tell when it's confidently wrong. Can't build the verification harness. Can't write the loop — because writing the loop *requires* knowing what good output even looks like, down at the code level. And there's data on the friction, too. One study of ten thousand developers found AI adoption was actually *lower* among senior engineers — partly out of well-earned skepticism.
Sam: Wait — so some of the exact people best positioned to run these fleets are the *slowest* to pick the tools up?
Alex: Which is its own quiet problem. So the senior's hands-on instinct turns out to be *necessary* — it's just not *sufficient* on its own.
Sam: So if it's not the pure senior, and it's not the pure architect… let me try to say it. The winner is the person who's *both*. Someone with the architect's judgment about what's even worth building — and the hands-on chops to catch the AI when it's lying.
Alex: Sam — that's the synthesis, almost word for word. The winner is the *hybrid orchestrator*. Architectural judgment about what's worth building, *and* hands-on fluency to know when the output is wrong and how to fix it. The old taxonomy — junior, senior, architect — is collapsing into one new role. And that role is exactly Cherny's own day. He's got the architectural judgment: he decides what the whole fleet of agents works on. And he's got the deep hands-on fluency: he built the harness, he reads the output. Both, at once.
Sam: So the theory was half right. Something *is* pulling ahead of the pure architect. It's just not the *unchanged* senior engineer.
Alex: That's the precise correction. The thing pulling ahead is the senior engineer who *became* an orchestrator. And — here's the unsettling tail — sometimes it's the domain expert. The accountant. Who skipped the engineering rung entirely.
Sam: That's the part that's going to keep me up. The accountant who can now build the thing — without ever having been a junior engineer at all.
Alex: And that brings us to the study I promised you — because everything so far has a believer's tilt to it, and a story that only quotes the believers is just marketing. So here's the finding that should puncture any "AI ten-x's every engineer" reflex. And it comes from a randomized controlled trial. The same design you'd use to test a drug.
Sam: A real RCT. On coding. Okay — set it up.
Alex: Mid-2025. A research group called METR takes sixteen experienced open-source developers. They have them do two hundred forty-six real tasks, on mature codebases the developers knew intimately — an average of *five years* on that code. Half the tasks, AI tools allowed — mostly Cursor, with the frontier Claude models of the day. Half, no AI. And before they start, the developers predict the AI will make them about twenty-four percent faster.
Sam: Sounds reasonable. Twenty-four percent. So what happened?
Alex: Afterward, they estimate it *had* made them about twenty percent faster.
Sam: Right, a little under what they hoped, but still —
Alex: It made them nineteen percent **slower**.
Sam: …Say that again.
Alex: Nineteen percent slower. *With* the AI. On the code they knew best.
Sam: But they *thought* they were twenty percent faster. So they were off by — what, nearly forty points? In the *wrong* direction? On their own work?
Alex: Every single person in the study was wrong, in the same direction, about their own productivity. They felt faster, and they were measurably slower.
Sam: That is genuinely unsettling. That's not "the tool underperformed." That's "I can't trust my own *sense* of whether the tool is helping me."
Alex: And that gap — between how it *felt* and what it *was* — is worth sitting on, because it's the most human part of the whole thing. Think about why it feels faster. When you hand a task to the model, *you* get to stop. You're not grinding through the boilerplate anymore — you fire off the prompt, you lean back, you wait, you skim what comes out. The effort on your end drops. And our brains read "less effort" as "more speed," almost automatically.
Sam: Oh — so the relief *is* the illusion. It feels easier, so I assume it's faster, even when the clock says I spent more total time prompting, waiting, reading, and fixing.
Alex: Exactly. It's like the difference between driving yourself and giving someone directions from the passenger seat. Driving yourself is more effort but you take the direct route. Directing someone else feels easier — you're just talking — but you're now describing every turn, catching their wrong moves, re-explaining the tricky bit. It can absolutely take longer, and the whole time it *feels* lighter, so you'd swear it was quicker.
Sam: And you'd defend that to anyone who asked. That's the scary bit — the people in the study weren't lying. They genuinely believed it.
Alex: They genuinely believed it. Which is exactly why you can't judge a tool like this on feel. You have to actually measure. And that's exactly why it matters — but here's the crucial part: it's not a refutation of the tools. It's a precise *map* of where their value is, and isn't. The best explanation: an expert working on familiar code is *already* operating near their ceiling. They know exactly what to type. So prompting, and waiting, and reading, and correcting the model — for them, that's pure overhead. Pure tax.
Sam: Whereas if I'm out of my depth —
Alex: — that's precisely when it helps. An unfamiliar codebase. A language you don't know. A domain you understand but can't yet express in code. And *that* — that's the mechanism behind Cherny's accountant. The domain expert gains enormously, because the alternative for them was *not being able to build it at all.* The expert engineer on home turf gains little or nothing, because they could already do it.
Sam: So the same tool is a rocket for the person who couldn't do it before, and a speed bump for the person who already could. The benefit is biggest exactly where the existing skill is lowest.
Alex: And there's an organizational version that's just as sobering. One analysis found AI lifted *individual* pull-request output by ninety-eight percent —
Sam: Ninety-eight — so basically doubled what one person pushes out.
Alex: Doubled individual output — and produced *no* statistically significant gain in *org-level* productivity. The whole system seems to hit a ceiling around ten percent.
Sam: How does that even happen? How do you double what everyone produces and the company barely moves?
Alex: Because the bottleneck just *moved downstream*. Picture a factory line. There's one slow station — say, packing the boxes — and that station sets the speed of the whole line, no matter how fast everything upstream runs. Now you supercharge the station *before* it — parts come flying out twice as fast. What happens?
Sam: They just… pile up in front of the slow station. The line doesn't go any faster. There's just a bigger heap of half-finished stuff waiting.
Alex: That's it exactly. In engineering, the AI supercharged the "writing code" station. But the slow station downstream — *reviewing* it, *testing* it, making sure it's safe and correct, merging it without breaking everything — that one didn't get faster. It got *more* loaded. Review time goes up, bug rates tick up, and the work piles up right in front of it. You didn't remove the constraint. You *relocated* it — and possibly made it worse, because now there's twice as much to inspect. So the honest multiplier — across the board — isn't ten-x. It's something like one-and-a-half to three-and-a-half-x for most people. Five-to-ten-x for the roughly one-in-five who are genuinely fluent in this new mode. And *negative* for an expert on code they already own.
Sam: So when someone tells you "AI made our engineers ten times more productive" —
Alex: — ask them to show you the org-level number, not the individual one. The gap between those two is where the truth lives.
Sam: Okay. We keep circling the juniors, and you keep saying that half is "strongly supported." I want the grim part now. What is actually happening to entry-level engineers?
Alex: This is where the evidence is the least ambiguous, and honestly the most consequential. Stanford ran a 2026 analysis of the labor data. Software-development employment for twenty-two-to-twenty-five-year-olds is down roughly *twenty percent* from 2024.
Sam: Twenty percent of the youngest engineers — just gone.
Alex: And here's the kicker that tells you it's AI and not just a soft economy: in the *same* AI-exposed roles, at the same kinds of firms, employment for workers over thirty actually *rose*.
Sam: Oh — that's the dagger. It's not "tech is down." It's specifically the *bottom* being cut while the top grows.
Alex: Specifically the bottom. And other readings of big datasets point the same way — junior employment falling within a couple of quarters of a firm adopting generative AI, internships shrinking, the junior share of engineering teams compressing. And the hiring managers say the quiet part out loud: a large majority report AI can already do intern-level work — and a majority say they *trust the model's output over an intern's.*
Sam: They trust the model over the human intern. That's brutal — but, okay, I almost get the logic. If the model does the intern's tasks faster and you trust it more, why hire the intern?
Alex: And that question — that totally reasonable question — is the trap. Because here's the mechanism, and it's the load-bearing risk in this entire transition. Yes: junior work is the boilerplate, the small bug fixes, the test scaffolding, the well-specified tickets. And that's *exactly* the work these agents do best — so it's the first thing automated. But —
Sam: — but that work wasn't *only* work.
Alex: That work was the *training pipeline*. It is *how a junior becomes a senior.* It's how judgment, and taste, and the ability to smell a bad design get built — over years of doing the unglamorous thing.
Sam: Oh. Oh, that's the problem, isn't it. You automate the bottom rung, and it *looks* like you just saved money —
Alex: — but what you've actually done is dismantle the machine that *produces the orchestrators of the future.* The exact hybrid people we just spent ten minutes calling the winners — they get *built* by climbing a ladder whose bottom rungs you're now sawing off.
Sam: So the bill doesn't come now. It comes later.
Alex: It comes in the 2030s — as a projected shortage of senior talent, when the cohort that should have been climbing finds the lowest rungs gone. And there's a cautionary tale already on the record. Klarna — the fintech — cut staff aggressively in favor of AI, publicly declared the era of big engineering teams over —
Sam: I remember this. Didn't they walk it back?
Alex: They walked it back. Quality suffered, and they quietly rehired. And the lesson there isn't "the automation was fake." The automation was real. The lesson is that the bottom rung was doing load-bearing work that nobody put on the balance sheet.
Sam: That's the line, isn't it. The bottom rung was load-bearing, and nobody costed it.
Alex: And speaking of things nobody costed — let's follow the money. Because none of this is free, and the economics might be the thing that actually bends the next two years.
Sam: Right — those "hundreds of agents," "tens of thousands of agents" — those aren't free to run.
Alex: They are emphatically not metaphors, and they are not free. Each one burns *tokens* — the units of AI computation you pay for. And at the frontier, a single fifty-turn agentic session can chew through something like a *million* input tokens.
Sam: A million. For one session. One task.
Alex: So you get a cost structure that looks nothing like classic software. Now — to be fair, the revenue side is staggering. Anthropic disclosed numbers around its thirty-billion-dollar Series G, in February 2026 — a round that valued the company at three hundred and eighty *billion* post-money. And they put Claude Code's run-rate revenue *above two and a half billion dollars*. More than doubled since the start of that year. Enterprise now over half of it.
Sam: Okay, those are not "struggling startup" numbers. That's a juggernaut.
Alex: It is. But here's the catch sitting underneath them. That revenue rides on top of a margin profile the whole industry is visibly straining against — because the inference, the AI computation, is a real, *recurring* cost of every single use. It can run a meaningful fraction of revenue. And that pulls gross margins well below the eighty or ninety percent that made software such a magical business in the first place.
Sam: Right — classic software, you build it once and every extra copy is basically free. This isn't that. Every single run costs you again.
Alex: Every run costs again. And we actually pulled this exact thread apart in an earlier episode — "Your $200 AI Plan Costs the Maker $14,000" — that flat monthly subscription that quietly costs the *provider* many multiples of what even a heavy user pays. Same dynamic here, at industrial scale.
Sam: So where's the strain actually showing up? Because "margins are thinner" is abstract.
Alex: It's showing up at the buyers, concretely. Uber reportedly burned through its *entire annual* AI-coding-tools budget — in four months.
Sam: Four months into a twelve-month budget. That's a fire.
Alex: Microsoft, by multiple reports, scaled *back* its Claude Code usage over soaring operating expense, and redirected toward a cheaper alternative. And a heavy individual user on a flat subscription can consume far more in raw inference value than they pay — which is exactly why the pricing's been so turbulent. New rate limits. Community complaints about how fast those limits get eaten. Brief experiments with pulling the tool from cheaper tiers — then reversing within a *day*.
Sam: So the forward implication is — what — that the real limit on all this isn't even the technology?
Alex: That's the underappreciated turn. For the next year or two, the constraint on AI-leveraged building might be less "can the model do it" — and more "can you *afford* to let it." The organizations that win won't just get the quality of their loops right. They'll get the *cost* right — observable, bounded, pointed only at work that justifies the spend.
Sam: So "write loops" comes with a silent second half: "write loops you can afford to run."
Alex: Write loops you can afford to run. That's the part the productivity euphoria leaves out.
Sam: Okay. Bring it home. One-to-two years out. Honestly — where is this trending, without the hype and without the doom?
Alex: So let me name the hinges, because that's the honest way to do a forecast. First, the capability curve is steep and real. By one well-cited measure, the *length* of task an AI agent can complete autonomously has been roughly *doubling every several months* — and the doubling has, if anything, been accelerating.
Sam: Unpack "length of task" for me — length how?
Alex: How long a stretch of work it can carry on its own before a human has to step in. A couple of years ago, that was measured in *seconds* — autocomplete, finish-this-line, a tiny snippet. Then it was a few minutes — write this whole function. Then a self-contained task that'd take a person half an hour. And the key thing is the *shape* of that growth: it's not adding a fixed amount each time. It's *doubling*.
Sam: And doubling is the one our brains are famously terrible at. The rice-on-the-chessboard thing.
Alex: That's exactly the right instinct. Doubling looks gentle, gentle, gentle — and then it goes vertical. If the length of task an agent can handle keeps doubling every several months, you cross from "minutes" through "a couple of hours" to "a full day of engineering work" startlingly fast. And the reason that matters for *this* story is that it's the difference between era two and era three. When an agent could only hold a task for a few minutes, you had to babysit it — that's the parallel-Claudes, context-switching world. The moment it can run reliably for an hour or more on its own —
Sam: — you can wrap it in a loop and walk away. With your laptop shut.
Alex: That's the unlock. The loops weren't possible until the task length got long enough to trust unattended. So when Cherny says "loops are as big a step as agents were," what made *loops* a thing you could even do is that curve. And that doubling is exactly what makes them viable *now* — and what made them science fiction two years ago. Extrapolate it, and you're looking at agents handling multi-hour, and eventually multi-day, engineering tasks within a couple of years. Anthropic's leadership talks in those terms. Even outside observers concede that fully self-written software is "conceivable within about two years." And Cherny's at the optimistic pole — "coding is solved for the kinds of coding that I do," he says — and he pictures "a hundred times more people writing code using agents," as the title "software engineer" dissolves into just "builder."
Sam: A hundred times more people building software. So that's not *fewer* coders — it's a flood of them, just… a different kind.
Alex: That's his vision. *But* — the same evidence names three hinges that vision turns on, and they're why the next two years will be messier than the slogans. Hinge one: the *eighty percent problem.*
Sam: What's the eighty percent problem?
Alex: Studies of agent runs find they reliably nail the visible eighty percent of a task — and miss the invisible twenty. The auth middleware. The cross-repo dependency. The migration. The stuff that lives *outside* the model's context window. And here's the subtle, crucial part: that failure is a *context-and-infrastructure* problem — not a "the model's too dumb" problem.
Sam: Pull those apart for me, because they sound the same from the outside. The agent failed either way.
Alex: They're really different, and the difference decides the whole forecast. "Too dumb" means: even with everything in front of it, the model just can't reason well enough. *That* gets fixed by a smarter model, and you wait. But "context-and-infrastructure" means the model is plenty smart — it just *couldn't see the whole problem.* It didn't know about the other six files it had to touch, because nobody showed them to it. It's like a brilliant contractor who does flawless work on the one room you walked them through — but there's a load-bearing wall in the room next door they never saw, and they cheerfully knock a hole in it.
Sam: Ah — so it's not that the contractor's incompetent. It's that you only handed them the floor plan for one room.
Alex: Exactly. And that reframes what's actually hard. The bottleneck for the next couple of years isn't only "make the model smarter." It's "get the *whole* problem in front of it" — the dependencies, the history, the bits that live in someone's head. Multi-file refactors and legacy code are where agents struggle most precisely *because* that's where the hidden context is densest.
Sam: So the last twenty percent — the invisible, load-bearing, "but what about the thing nobody documented" twenty percent — is exactly the part the human still has to own.
Alex: Which, notice, is the *orchestrator's* job all over again. Hinge two: security and quality debt. A large share of AI-generated code ships with design flaws, or known vulnerabilities. Review time grows. And a meaningful fraction of enterprise agentic-AI projects are *forecast to be cancelled* by 2027, as the gap between the slick demo and actual production bites.
Sam: The demo-to-production gap. The graveyard of every cool technology.
Alex: And hinge three is the pincer we already walked through — the dismantled training pipeline, and the token tax. So put it all together, and the realistic one-to-two-year picture is neither utopia nor collapse. It's a profession reorganizing around a *smaller* number of higher-leverage orchestrators — each running fleets of agents against a verification harness. A durable premium on that hybrid skill: architectural judgment *plus* hands-on fluency. A real and widening gap between the one-in-five who are fluent in this mode and everyone else. A genuine economic constraint from inference cost, disciplining which loops are even worth running. And a quietly broken bottom rung, whose consequences land in the 2030s.
Sam: So the whole world starts to look like Boris Cherny's desk.
Alex: That's the image to leave with. The frontier looks like his desk — a human writing loops, a fleet of agents writing the code, and a verification harness deciding what's actually true — generalizing outward. From the people who build the models, to the domain experts who use them. And so the decisive question, for any one of us, is *not* "will AI replace me."
Sam: It's which person am I becoming.
Alex: It's: am I becoming the orchestrator who has *both* the judgment and the hands? Or am I clinging to the rung that's being automated — or to the editor that's being uninstalled?
Sam: Okay. That genuinely reframed it for me. Can we land the plane — what do we actually walk away with?
Alex: Let's recap it tight. One: the most credible witness to the future of engineering is the man who automated it — Boris Cherny — and his testimony is more unsettling than either side's slogans. He doesn't write code, he writes loops, and he says the biggest winners aren't even professional engineers. Two: the popular theory lands *half* right. Juniors really are being displaced — sharply, measurably — and the bottom rung that *makes* seniors is being sawn off in the process. But "senior beats architect" is a false binary. The winner is the hybrid orchestrator who fuses architectural judgment with hands-on fluency — and the role of "software engineer" isn't vanishing so much as moving up one level, to the person who writes the loops.
Sam: And three — the one I'll actually carry around — there's that randomized trial where the experts got *slower* and couldn't even feel it. So your own sense of whether AI is helping you is not to be trusted. Check the number, not the vibe.
Alex: Check the number, not the vibe. And the next two years come down to four hinges — how fast the models close that invisible-twenty-percent gap, whether the security and quality debt gets paid down, who can afford the token tax, and what happens to a generation that never got to climb the bottom rung. Watch *those* — not the headline that "AI writes all the code now."
Sam: Because the code was always the easy part.
Alex: The code was always the easy part. The judgment about what to build, and the verification of whether it's right — that's the job that's left. And it's the job that's *growing.* I really hope you came away from this seeing a little more clearly where all of this is heading. It's a genuinely complex, fast-moving picture, with a brutally short shelf life on what's true — and honestly, that's exactly what makes it worth following closely, instead of letting it just wash over you.
Sam: And before we go — a quick note on how the show's made, because we like to be straight with you.
Alex: For full transparency: this show is AI-generated. Dan builds a custom stack of AI tools to chase the questions he can't stop thinking about — it actually started out made with NotebookLM, and it's now produced with his own engine — mainly so he can learn all of this himself, and he publishes it for anyone who'd like to learn along with him.
Sam: Which feels exactly right for an episode about a man who automa…