Why OpenAI Killed Sora, and Google Owns AI Video

An episode of Dan's AI Intel

Rank image, video, and voice by compute cost and you get a perfect map of the AI race — and why only Google can afford all of it.

Published · By Dan Walter

Transcript

Alex: In the spring of 2026, a company with one of the most valuable products on Earth reached over and switched off an app that ten million people had downloaded.

Sam: Ten million — and they just killed it? Why would anyone do that?

Alex: Because the exact same chips running that app earned them a thousand times more money doing something else. And that one decision draws you a near-perfect map of who's winning all of AI.

Sam: Okay, a thousand times. That's not a business problem, that's a physics problem. I need the whole story.

Alex: Welcome back to Dan's AI Intel — the show where we take the one question that actually matters this week and dig underneath the headlines until it makes sense. I'm Alex, here as always with Sam.

Sam: Hello. And this is one of those episodes where I came in thinking I understood the picture, and I'm told I've got it about half right.

Alex: That's exactly the shape of it. So here's the trigger. OpenAI built a video app called Sora, it blew up, and then this year they quietly turned it off. And the tempting read is: everyone rushed into AI image and video a couple of years back, and now it's just Google and a sliver of OpenAI left standing.

Sam: Which is roughly the story I had in my head, yeah. The field thinned out.

Alex: Half of that is true and half is an optical illusion, and the thing that separates them is a single number — how much computer time it costs to make one second of what you're looking at. Rank voice, images and video by that number, and the whole market snaps into focus.

Sam: So the promise is: give me one number, and I can predict who owns each of these markets?

Alex: That's the promise. We'll follow it from a killed app, down into the actual mathematics of the transformer — why a picture is not a thousand words, why video's cost literally explodes, why voice is the cheap frontier everyone's ignoring — and then back up to what it all reveals about Google, and about where this whole thing is heading, which is somewhere much bigger than making clips.

Sam: And I want to know why it had to be Google specifically. Hold that thought.

Alex: I will. If you're finding the show useful, by the way, follow us wherever you're listening — it's free, and it's genuinely the biggest thing that helps a small independent show like this keep going. Right. Let's start with the switch-off. So picture the timeline. Late 2025, OpenAI launches Sora as a standalone consumer app. It's a hit — close to ten million downloads. And by spring 2026, they've discontinued it. Retired the app, wound down the interface, no replacement video product announced.

Sam: And this is a genuinely popular thing. Ten million downloads isn't a failed launch. So what did Altman actually say?

Alex: Altman's explanation was blunt. He said the company needed to concentrate its compute and its people on the next generation of systems, and video was a side quest it could no longer justify feeding.

Sam: "Feeding" is doing a lot of work in that sentence.

Alex: It's the whole sentence, really. Here's the arithmetic underneath it. Analysts covering the company estimated Sora was burning something on the order of fifteen million dollars a day in raw running costs — the compute to generate all those videos.

Sam: Fifteen million a day. Okay. Against what in revenue?

Alex: Against roughly two million dollars across the app's entire lifetime.

Sam: Wait. Say that again. Fifteen million a day going out, and two million total — not per day, total — coming in?

Alex: Total. Over months. Each ten-second clip cost about a dollar-thirty to make and sold, effectively, for nothing.

Sam: That's the thousand-to-one you opened with. That's not a leaky bucket. There's no bucket.

Alex: And I should be honest about the sourcing — those are external analyst estimates, not audited figures OpenAI published. But nobody inside the company disputed the order of magnitude, and the order of magnitude is the whole point. When you're wrong by a factor of a thousand, the exact decimal doesn't save you. And now the part that turns a sad story into a strategic one. Those same graphics chips, pointed instead at OpenAI's coding assistant and its enterprise business, were part of an operation run-rating around twenty-five billion dollars a year.

Sam: Ah. So it's not that the chips are unprofitable. It's that the chips have a much better job to go to.

Alex: Exactly right. This is the key move, and it's worth slowing down on, because it's the hinge the entire episode swings on. A chip isn't loyal to a product. It's a unit of compute, and it flows to wherever it earns the most.

Sam: So think of it like a taxi driver in a city. The driver doesn't care whether you want to go to the airport or round the block. They'll take the fare that pays best.

Alex: That's exactly it. And a chip writing code for a paying enterprise is the airport run — it earns a fortune. The very same chip making a free video is the trip round the block that doesn't even cover the petrol. There was even a reported Disney partnership, worth something like a billion dollars, that fell apart in the fallout — and faced with chips that lose money on video versus chips that print money on code, there wasn't really a decision to make.

Sam: So the shutdown isn't OpenAI saying "we're bad at video." It's them confessing that at the quality frontier, video doesn't pay for its own electricity.

Alex: That's the sentence. Hold onto it, because it's the thesis of the entire episode. Switching off Sora is a Rosetta Stone — it tells us that only a company that can either eat that loss, or escape it with cheaper chips, gets to stay at the video table at all.

Sam: Okay, but you've told me video is brutal and code is lucrative. What I don't have yet is why. Why is a video so much more expensive than a chatbot answer in the first place?

Alex: And before the "why," let me plant one more idea, because it's the reason this whole story is worth telling. Most arguments about who's ahead in AI happen in a fog — benchmarks get gamed, "vibes" are subjective, the models that matter most are often ones nobody outside a lab has ever touched.

Sam: Right, it's all vibes and leaderboards.

Alex: But generative media strips the fog away, because here the economics are naked. A chatbot answer is cheap enough that its cost barely constrains anyone's strategy. A minute of high-definition video is expensive enough that the cost dictates everything — who can offer it, at what price, to how many people, whether the product can survive contact with an accountant at all.

Sam: So it's a rare place where you can actually see the money.

Alex: And here's the principle underneath that. When a resource is abundant, watching how it's shared tells you almost nothing. When a resource is scarce, watching how it's rationed tells you exactly who holds the power. Compute is now the scarce resource — and image, video and voice are where you can literally watch it being rationed in real time.

Sam: Okay. That's the sell. So let's follow the rationing. Give me the number.

Alex: Right, and that's the number that runs the whole show. Let me set up the gradient before we go spelunking into the math, because the gradient alone explains almost everything.

Sam: Go.

Alex: Three ways an AI can make something a human perceives: a voice, a picture, a moving image. Rank them by how much compute a second of output costs. At the cheap end, voice. Generating a second of natural speech is only modestly pricier than generating text.

Sam: And "cheap end" means what, in market terms?

Alex: It means a well-funded startup can just win it outright — no giant required. In the middle, still images. Expensive enough that you need serious infrastructure, cheap enough that a dozen credible players can all afford to compete, and nobody runs away with it.

Sam: And then video at the far end, which we've established is a bonfire.

Alex: A bonfire. A few seconds of video can cost more than thousands of still images. And so only the largest, most vertically integrated companies on the planet can sustain the frontier. And here's the pattern that turns this into an actual law rather than a list. In every one of these markets, the cost of a single unit of output sets the ceiling on how many companies can afford to fight for it.

Sam: So the price of one output decides the size of the crowd. Cheap output, big crowd. Ruinous output, tiny crowd.

Alex: That's the whole law, in two sentences. Read the frontier from cheapest to costliest and you're reading the market from most open to most concentrated, in the same breath — voice, a startup owns it; images, a whole crowd competes; video, a handful of titans and nobody else.

Sam: I love a rule like that — but I'm also suspicious of it, because it sounds too tidy. So convince me. Why is the gradient this steep? Why is voice a trickle and video a flood? Because "video just has more pixels" feels too easy.

Alex: It is too easy, and it's wrong in a really instructive way. Which is where we have to look at what the machine is genuinely doing when it paints a picture. And almost everyone's mental model of that — including mine, until I really looked — is off. The intuitive story goes like this: an image is expensive because the model writes a little caption for every pixel. A million pixels, a million tiny descriptions, so obviously it costs a fortune.

Sam: Right, that's exactly what I'd have guessed. "A picture is worth a thousand words," so the computer's writing all thousand of them. Times a million.

Alex: And that's not how a modern image model works at all. The first thing a model like Stable Diffusion or Google's Imagen does is refuse to work in pixels.

Sam: Refuse how?

Alex: It has a separate little network that compresses the picture first. Take a two-fifty-six by two-fifty-six image, and it collapses down to a thirty-two by thirty-two field of numbers — roughly an eightfold shrink in each direction. The big model never sees raw pixels. It sees this compact, coded-down grid, chopped into patches, and it treats each patch as a token — exactly the way a word is a token in a sentence.

Sam: So the model isn't staring at a million pixels. It's looking at a few thousand tokens. That's a big deal for the "million captions" story, because there's no million anything.

Alex: There's no million anything. A still image might be only a few thousand tokens. So far, this is cheap. The expense doesn't come from the size of the picture at all. It comes from the second idea, which is called diffusion. Instead of writing the picture in one shot, the model starts from pure static — literal visual noise, like an old TV tuned to a dead channel — and it removes a tiny bit of noise at a time. Each pass, it asks itself, "what would this look like if it were just slightly less noisy?" And it repeats that pass. Twenty times, fifty, sometimes a hundred, before the image finally resolves.

Sam: So it's not painting, it's... developing a photo. Slowly bringing it out of the fog, over and over.

Alex: That's a lovely way to put it, and it's exactly the right instinct. And here's the crucial contrast with text, because this is where the cost hides. When a language model writes a sentence, it produces each word in a single forward pass — and thanks to a trick called caching, it never has to redo the work for words it's already written. It writes a word, files it away, and never looks back.

Sam: Whereas the image model...

Alex: The image model reprocesses its entire grid of tokens, dozens of times over. Every one of those denoising passes chews through the whole picture again, start to finish.

Sam: Oh. So the cost isn't the number of pixels. It's the number of times it re-reads the whole thing.

Alex: Exactly — the cost isn't the pixels, it's the re-reads. And when you actually add it up, one picture from a model like that costs on the order of two hundred trillion operations.

Sam: Two hundred trillion. I genuinely can't feel a number that big. Give me a comparison.

Alex: That's roughly the compute of generating a full page of dense text — squashed into a single image. So a picture isn't worth a thousand words.

Sam: It's worth a thousand re-readings of the same paragraph.

Alex: That's the line. And hold that phrase — the entire grid, reprocessed dozens of times — because the second you add the dimension of time, that phrase detonates.

Sam: Okay, I can feel where this is going, and I'm slightly scared of it. Take me into video.

Alex: So a video isn't one image. It's many images that also have to agree with each other from frame to frame — the coffee cup in the corner has to stay a coffee cup as the camera moves, the light has to fall the same way, the person can't grow a third arm between frames.

Sam: And I'm guessing they don't just make each frame separately, like a flipbook.

Alex: They don't, and the reason why is the whole ballgame. If you generated each frame in isolation, they'd never agree — you'd get a flickering mess. So a model like Sora carves the clip into what are called spacetime patches — little bricks of video, each one covering a small square of the screen across a short slice of time. And it feeds those bricks in as tokens, exactly like the image patches. There are just vastly, vastly more of them.

Sam: How many more are we talking?

Alex: Five seconds of fairly ordinary video works out to more than eighty thousand tokens.

Sam: Eighty thousand — for five seconds. To put that in human terms: a page of text is about a thousand tokens. So five seconds of video is a longer read, for the machine, than a whole book chapter.

Alex: Longer than a book chapter. And that alone would just be "large." It becomes ruinous because of the one feature of the transformer that's magical and cursed at the very same time. It's called attention. Attention is how the model keeps a scene coherent. Every token in the sequence looks at every other token — that's literally how it makes the cup stay a cup while everything around it moves.

Sam: Every token looks at every other token. Okay. That already sounds expensive.

Alex: Here's the killer. Because every token has to look at every other one, the cost doesn't grow with the number of tokens. It grows with the square of the number of tokens.

Sam: The square. So if I double the length of my video...

Alex: You don't double the cost. You roughly quadruple it.

Sam: Let me make sure I actually feel that. It's like a dinner party where every guest has to shake hands with every other guest. Add a few more guests, and the number of handshakes doesn't creep up — it balloons.

Alex: That is the perfect analogy — the handshake problem. Ten guests is forty-five handshakes; twenty guests isn't ninety, it's a hundred and ninety. Double the guest list, roughly quadruple the handshaking. And with video, you're inviting eighty thousand guests, all shaking hands, on every single denoising pass.

Sam: And I'm guessing pushing the resolution up does the same thing.

Alex: The same curse. Sharper picture, more tokens, and the squared cost punishes every one of them. Researchers measured that training on a longer, higher-resolution clip ran forty times slower than a short low-res one — from that one squaring term alone.

Sam: Forty times. From geometry, basically. Not from a worse model — just from the shape of the math.

Alex: Just the shape of the math. And in a video model, that attention step alone can eat more than eighty-five percent of the total compute. So now stack the whole thing up: way more tokens to begin with, times the square of those tokens for attention, times the dozens of denoising passes that diffusion demands. And you land on the number that killed Sora — a minute of high-quality video costs roughly what eight hundred to fifteen hundred still images cost.

Sam: So video isn't one rung above images on the ladder. It's three multiplications above them.

Alex: Three multiplications above. More tokens, times their square, times the repeated denoising. And this is the deepest answer to your original "why." Video isn't token-hungry by a bit. It's token-hungry by a compounding law of arithmetic that no amount of clever engineering has managed to repeal yet.

Sam: Right — so we've done images as a thousand re-readings, and video as that same idea squared and then hammered by diffusion. Which sets up the total opposite case. You keep teasing that voice is cheap. If video is the flood, what makes voice the trickle?

Alex: So run the exact same lens over voice, and it flips completely. Audio gets turned into tokens too — but by a neural codec that is astonishingly stingy.

Sam: Stingy how? Give me the token count so I can hold it against that eighty thousand.

Alex: Where a video needs tens of thousands of tokens per second, high-quality speech needs only around fifty to a hundred tokens per second. A full minute of talk is maybe fifteen hundred to two thousand tokens.

Sam: Hang on. A whole minute of speech is fewer tokens than one second of video?

Alex: Smaller than a single second of video. And it gets better — remember diffusion, the dozens of passes over a giant grid? Voice doesn't do that. Speech is generated the cheap way, one token after another, like text, with that same filing-it-away trick. So computationally, voice is the closest cousin of text in the whole generative family.

Sam: So that's why the big labs never went to war over it. There's no giant compute moat to defend, so it didn't attract the players whose whole strategy is outspending everyone.

Alex: Right — no ruinous compute moat means nobody had to be a titan to compete, so the titans never bothered turning up. Where video punishes everyone who isn't huge, voice barely charges an entry fee. And so, instead of the giants, voice got won by specialists. The headline name is ElevenLabs — a company most people outside the field have genuinely never heard of. They built the best-regarded synthetic voices in the world, and by mid-2026 they were in talks at a valuation around twenty-two billion dollars.

Sam: Twenty-two billion, for the voice company nobody's heard of. That's not a niche.

Alex: And roughly double what they were worth just five months earlier — so it's accelerating, not settling. Others — Cartesia, Deepgram — took the low-latency and transcription corners. The whole text-to-speech market cleared north of six billion dollars a year, and it doubled in three years.

Sam: And yet it's the quiet one. Which feels like the setup for you telling me that's a mistake.

Alex: It's a strategic blind spot, and here's the argument. Voice isn't a novelty layer you bolt on at the end. It's the interface to every AI agent that's ever going to hold a conversation, take a phone call, or read the world to you while you're driving.

Sam: So when you picture the future where you're just talking to your AI all day — the thing doing the actual talking is this cheap, ignored layer.

Alex: That's the whole point. The cheapest thing to compute turns out to be one of the most valuable things to own — because it's the doorway, not the room.

Sam: The doorway, not the room. Say more, because that's the bit that flips it for me.

Alex: The room is the reasoning, the intelligence, the expensive stuff everyone's fighting over. But you never get into the room without walking through the door — and for an agent, the door is its voice. Own the doorway and you sit between the human and every clever thing behind it. Quick aside, actually — if that "the cheap thing turns out to be the strategic thing" twist grabs you, we did a whole episode on the same shape hiding inside Google's economics: the hidden seventy-times subsidy buried in its two-hundred-dollar plan — that's number 20, from a few months back. Same species of surprise. Anyway — voice is settled. Which brings us back to the flood. Video. And the puzzle we opened with.

Sam: Right, because you told me the "only Google and OpenAI left" story was half an illusion. So where's the actual competition hiding?

Alex: So take the compute logic we just built, point it back at video, and the apparent collapse of competition dissolves into something much more interesting. It's not that the contest stopped. It's that the contest moved.

Sam: Moved where?

Alex: Rank the best text-to-video models in 2026 by blind human preference — people picking which clip looks better, without knowing who made it — and the top of that board is almost entirely Chinese.

Sam: Chinese. Not Silicon Valley.

Alex: Number one is Kling, made by Kuaishou. ByteDance — the company that owns TikTok — has Seedance. Alibaba's got Wan, and a newer one called HappyHorse. MiniMax has Hailuo. These aren't lab experiments; they're shipping products from firms that already run the biggest short-video platforms on the planet.

Sam: And the American names?

Alex: Among Western companies, Google's Veo sits near the top as very nearly the only non-Chinese model at the frontier. And here's the one that really tells the story — Runway, the American startup that actually led this whole field at the end of 2025, has slipped out of the top ten entirely.

Sam: So the story isn't "the West won and everyone else gave up." It's almost the reverse. The West mostly left, and China's at the top. Which really breaks my mental model, because I'd have assumed a frontier this cutting-edge would be a Silicon Valley thing. But why China? What is it about those specific companies?

Alex: Because the compute logic selects for a very particular kind of company — and it is not "the one with the best researchers." To sustain a video model you need two things in bulk: enormous cheap compute, and a bottomless supply of video to train on.

Sam: And a company like ByteDance or Kuaishou already has both — because that's just their day job. They already run TikTok-scale video for billions of people.

Alex: That's the entire point, and it's the whole asymmetry. The marginal cost of also running a video model is something they can bury inside an infrastructure budget they were already paying anyway — the chips are already spinning, the video is already flowing. Whereas American AI is dominated by pure-play labs that rent their compute at a premium and don't own a video platform of their own, so for them every one of those costs is a new bill.

Sam: So the frontier didn't reward the best model-builders. It rewarded the companies whose existing business already happened to pay for the single hardest input.

Alex: That's the whole thing in one line. The day job paid for the moat. And quick aside — if that "China quietly at the frontier" theme is pulling at you, we went deep on exactly that surprise a few weeks back, in our episode on China hitting the AI frontier and then giving it away — number 33. Worth a listen alongside this one. But it sets up the obvious next question, doesn't it —

Sam: It does, and I want to ask it straight. If this game rewards owning a video platform's worth of chips and data, and that mostly describes Chinese giants — how is Google the one Western company that gets to sit at the table? What does Google have that OpenAI doesn't?

Alex: Right. So Google is the singular Western case, and the reason is almost embarrassingly clean. It's the one company outside China that owns all three of the things the compute era rewards — and it owns them outright, in one building.

Sam: Three things. Walk me through them.

Alex: Start with the chips. Google designs its own AI processors — they're called TPUs — and runs them in its own data centres.

Sam: Whereas everyone else is buying from Nvidia.

Alex: Buying, or really renting, from Nvidia, at a steep markup. There's a thing the industry half-jokingly calls the "Nvidia tax" — the gap between what a top chip costs to actually manufacture, a few thousand dollars, and the twenty to thirty-odd thousand dollars a buyer pays for it on the open market.

Sam: So the chip that costs a few thousand to build sells for ten times that, and everyone who isn't Nvidia just... pays it.

Alex: Everyone who has to buy on the open market pays it. Google mostly doesn't, because it makes its own. And when you skip that tax, the numbers get dramatic. By credible estimates, Google's TPUs deliver something like four times the useful work per dollar that rented GPUs do — call it an eighty-percent cost edge on the workload that eats most AI compute.

Sam: Okay, but let me stress-test that. An eighty-percent cost edge sounds huge on a spreadsheet — but does it actually change who can play? Or is it just Google having fatter margins?

Alex: It's the difference between a live product and a dead one — and here's exactly why. When your output is video, and video is a compute bonfire, a four-times cost advantage isn't a rounding error on the margin. It's the line between a product that's merely very expensive and one that's flatly impossible. OpenAI hit "impossible" and switched Sora off. Google, running the same kind of workload at a quarter of the cost, could keep the lights on.

Sam: So the same fire that burned OpenAI's video app down is a fire Google can afford to stand right next to — because its fuel costs a quarter as much.

Alex: That's it exactly. Same fire, quarter-price fuel. And the more expensive the workload, the more that quarter-price fuel matters — which is why the edge shows up most starkly in video of all things.

Sam: Right — so that's the chips. You said three things. What are the other two?

Alex: Then there's the data. Google owns YouTube. And it's confirmed that it trains its models on a corpus of some twenty billion videos — the largest, most varied library of moving images that has ever existed, already labelled by titles and captions and by what humans actually chose to watch, sitting inside the same company that builds the model.

Sam: Twenty billion videos. And nobody can just go and buy an equivalent, because there isn't one to buy.

Alex: There's no second YouTube on the shelf. You can't acquire your way to it; there's only one, and Google owns it. And then the third thing — distribution. Search, Android, Chrome, Workspace, and YouTube itself put whatever Google generates in front of billions of people at basically no cost to acquire them.

Sam: So let me put the three together. Its own chips, so the bonfire is cheap. Its own video, so it has the fuel to train on. And its own pipes, so it doesn't have to pay a cent to go find an audience.

Alex: Chips, data, and the pipes — the three scarce inputs of the entire compute era, all under one roof. That's why Google, uniquely in the West, can afford to field a frontier model in images, in video, and just fold generation straight into its assistant — instead of being forced, like OpenAI was, to pick one and kill the rest.

Sam: And that's the resolution of the puzzle we opened with. The "diversity is gone" feeling is real — but only at the expensive end. And it's not that competition died. It's that a compute filter let through exactly the companies that already owned the stack.

Alex: And this is where it stops being a story about a creative-tools market, and becomes something a lot bigger. Because video generation is quietly turning into something else entirely.

Sam: Into what?

Alex: Think about what a video model fundamentally does: it predicts the next frame. Push that same machinery far enough, and it can predict the next frame of an interactive world — one that responds to a controller, obeys a rough physics, and stays coherent as you move around inside it.

Sam: So not "generate me a clip of a street." More like "generate me a street I can walk down — and it keeps making sense as I go."

Alex: Precisely that. And Google's DeepMind has already shown it working — models called Genie that generate navigable worlds in real time, running at around twenty-four frames a second, holding together for minutes at a stretch, that a person, or more importantly an AI agent, can just explore. A world no one actually built, being dreamed up frame by frame as you move through it.

Sam: Okay, why does "an AI agent can explore it" matter more than a person exploring it? That's the bit I want.

Alex: Because a world you can act inside is a training ground. It's a place to teach a robot to grasp and to walk, or to teach an agent to plan — cheaply, safely, a million times over, without smashing real hardware in the real world.

Sam: Ah — so instead of buying a thousand real robots and breaking most of them learning, you spin up a million simulated ones overnight.

Alex: A million simulated ones overnight, and if they fall over, nothing actually breaks. So the frontier of "make me a video" is quietly merging with the frontier of general intelligence itself. The toy becomes the training ground for the real thing.

Sam: And — let me guess — it's merging on the turf of the one Western company that already owns the chips to run it, the video to train it, and the reach to ship it.

Alex: On Google's turf. Which is where the whole board finally becomes legible. The apparent diversity is real — but it lives at the cheap end, in voice where startups thrive, and in images where a dozen strong models compete. The apparent concentration is real too — but it was never a lack of competition. It was a compute filter. And at the costly end of that filter, the survivors are China's platform giants and, alone in the West, Google.

Sam: Compute intensity was destiny the entire time. It decided who could afford each frontier, it decided the video race would be fought mostly outside America, and now it's deciding who gets to build the simulated worlds that come next.

Alex: Right — and that's the law the whole episode was walking toward. The landscape looks chaotic right up until you weigh every output by the compute it burns — and then it snaps into one clean line: the cost of making a thing decides how many companies can afford to make it. So let's land it. If you take three things out of today, take these.

Sam: One: the map. Voice is cheap, so it belongs to startups. Images are moderate, so they belong to a crowd. Video is ruinous, so it belongs to the tiny handful of firms that own their chips, their data and their distribution — which, outside China, is exactly one company.

Alex: Two: the machine underneath. A picture is expensive not because it has a lot of pixels, but because the model re-reads its whole grid dozens of times. Video takes that and squares it — the handshake problem at eighty thousand guests — which is the real reason a minute of it costs what a thousand images cost, and the real reason OpenAI walked away.

Sam: And three: the "so what." Switching off Sora wasn't a retreat from a bad product. It was an honest confession that at the frontier, video doesn't yet pay for its own electricity — and the companies that can eat that cost are the ones for whom the electricity was already sunk.

Alex: And the question worth actually watching isn't who has the best model next year. It's whether anyone ever finds a way to break that quadratic arithmetic — the squared cost that makes video so brutal. Because until someone does, the frontier keeps belonging to the owners of the stack. And the most important of those owners is Google.

Sam: That's the one that reframes it for me. We spend all our time arguing about which model is smartest, and the real gatekeeper turns out to be an electricity bill.

Alex: That's the show. And honestly — thank you so much for spending this time with us. If you came away seeing a bit more clearly where all of this is heading, then it did its job: it's a genuinely fast, complicated, high-stakes picture, with a brutally short shelf life on what you know — and that's exactly what makes it worth following closely.

Sam: One honest note on how this show is made, too.

Alex: It's AI-generated. Dan builds a custom stack of AI tools to research, analyse, verify and illustrate the questions worth understanding — mostly to learn them himself, and he publishes it for anyone who'd like to follow along. AI-assisted, fact-checked, and always worth a second look.

Sam: And before you go, one genuinely useful thing you can do.

Alex: Follow the show. Whatever app you're listening in right now, there's a follow or a plus button — one tap, it's free, and it does two real things. You'll get every new episode the moment it lands, and honestly, for a small independent show like this, a follow is the single biggest lever there is for helping it reach other people trying to make sense of all this. So if today was worth your time — go ahead and hit follow.

Sam: And one last thing, because it genuinely shapes what we do. If there's a claim in here you'd push back on, or a thread you want us to pull harder on next time, tell us — the address is podcast at connectiveshift dot com. We read every single message, and it really does decide what we dig into next.

Alex: So — what should we weigh by its compute cost next? Tell us. Until then, thanks for listening, and we'll see you in the next one.