AGI Check-In: The Benchmarks Broke. Science Didn’t.

An episode of Dan's AI Intel

The honest measure of AI progress is no longer a score — it's what survives a check the model can't influence.

Published · By Dan Walter

Executive summary

In the first week of September 2026, the same model scored 99.9% and 62.7% on the same test. GPT-6 Astra's headline result on ARC-AGI-3 came from a harness OpenAI built; the independent scorer, running its own rig, published a number thirty-seven points lower — and the cheaper run was the one that scored higher. In the same week, 354 proteins that no human designed came back from a wet lab in Switzerland having bound to their targets. One of those results is a claim about a model. The other is a fact about the world, and no amount of prompting could have produced it.

That gap is the answer to the question everyone is actually asking. Benchmark scores stopped being usable instruments this generation: the leading independent ruler for agent capability now says its own measurements above sixteen hours are unreliable, and the two flagship models each win on the benchmark the other one chose. What replaced them is narrower, slower and far more trustworthy — the rate at which model output survives a check the model cannot influence. A proof assistant that does not care how confident the prose sounds. A computer algebra system. A mass spectrometer. The mathematician whose conjecture it is. On that ledger the acceleration is real and it is recent: the score on running an experiment end to end roughly doubled in twelve weeks, and the share of protein designs that actually bind went from roughly a quarter to roughly half.

It is also narrower than the announcements, and it has moved somewhere the public cannot follow. Where no independent check exists — a molecule's behaviour in a human body — the record has not moved at all: still zero approvals for a drug whose target and molecule were both found by AI, and second-phase trial success indistinguishable from the historical rate. And the most capable versions of both flagship models are now formally withheld. Anthropic's Mythos 5.1 is Fable 5.1 with the safeguards taken off, released only to vetted defenders and life-science labs. OpenAI's life-sciences model is available in one country to a handful of named companies. The suspicion that the labs keep their best work back is no longer a suspicion; it is printed on the box, and it means the public scoreboard and the actual frontier have quietly stopped being the same thing.

Three labs emptied the shelf in seventy-two hours

Between the first and the third of September 2026, the entire frontier moved. Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1 on the 1st. Google released Gemini 3.8 Flash on the 2nd, alongside a locked-down security sibling called 3.8 Flash Cyber. OpenAI began rolling out GPT-6 Astra on the 3rd. Three labs, three days, every flagship refreshed at once — and at the Astra briefing OpenAI's president Greg Brockman said he personally believes the company has reached artificial general intelligence, then closed with the line: "Welcome to the AGI era."

That is a remarkable sentence to say out loud, and it lands on a question this show keeps circling: how fast are we actually moving, and how would anyone outside these companies know? The instinct is to look at the launch tables. That instinct is now broken, and understanding why it broke is more useful than any single number on them.

Because something else happened in the same fortnight, quieter and much harder to fake. For the first time, a meaningful volume of what these models produced was checked by things that are not models — a proof assistant, a computer algebra package, a mass spectrometer, a plate of engineered proteins in a laboratory freezer. Those checks are slow, unglamorous, and completely indifferent to how fluent the output sounded. They are also the only measuring stick left that a marketing department cannot tune. So this check-in measures the AGI curve the way you would measure any other scientific claim: not by what was announced, but by what survived being checked.

What actually came out, and who checked it

Start with the scan, because the raw list is genuinely startling and most of it is real.

In mathematics, five named open problems moved in four months. An OpenAI reasoning model disproved Erdős's 1946 unit distance conjecture, announced in May; nine mathematicians including Fields medallist Timothy Gowers then produced a human-verified version of the counterexample. In July, Claude Fable 5 produced an explicit counterexample to the Jacobian conjecture in three variables at degree 7 — presented by Levent Alpöge, digested publicly by Terence Tao two days later, and checkable in any computer algebra system in seconds. GPT-5.6 Sol Ultra, running 64 subagents in parallel, generated a proof of the fifty-year-old Cycle Double Cover conjecture in under an hour; it has since been formalised in Lean 4 several times over by separate groups. A neurosurgery resident in Beijing ran GPT-5.6 Sol for about sixteen hours and came out with a proof of Crouzeix's conjecture, open since 2004 — and Michel Crouzeix checked it himself and thinks it is correct. And Astra improved a bound on short gaps between primes from 240 down to 186, with a Lean formalisation attached.

In protein design, the results left the screen entirely. Across fifteen biological targets, Anthropic's models produced 1,320 candidate binders; Adaptyv Bio and Twist Bioscience synthesised and tested them, and 354 of them bound, hitting 14 of the 15 targets. Hit rates ran from 22.6% to 35.1% depending on configuration, against an industry norm of 10 to 15%. On one target, RBX1, the model hit 40% where the human entrants in a public design competition managed 3.7%. With Mythos 5.1 the September figure is close to 50% across twelve targets, and on three of them the binding affinities came in ten times higher than the best designs any human submitted.

In analytical chemistry, Claude was handed raw instrument files in undocumented vendor formats, worked out the encoding itself, and returned an NMR structure determination in 23 minutes and an LC-MS purity analysis in 19 — landing within 0.08 of a hydrogen on the atom count and calling the purity at 96.4% against the contract laboratory's 96.33%.

In planetary science, Claude trained a neural network on radar data NASA's Magellan orbiter collected more than thirty years ago and produced a new elevation map of a third of Venus, resolving features at two to three kilometres instead of ten to twenty and improving height accuracy by up to 25%. Anthropic released it under a Creative Commons licence specifically so that NASA's VERITAS and ESA's EnVision missions can use it to choose what to photograph when they arrive.

And in computation itself, Mythos 5.1 rewrote the GPU kernels of seven open-source deep-learning models, hitting speedups of up to 2.5× with bit-identical outputs and cutting the cost of genome-wide analyses by 30 to 60% — work Anthropic says would normally occupy a team of performance engineers for weeks.

Six domains. Every one of those results is public. But they are not equally true, and the difference between them is not how impressive they sound.

Exhibit — The results that matter are the ones something other than a model agreed with. Read down the last column, not the first: it is the only one a press release cannot fill in. Source: Anthropic model card and research notes (Sep 1 2026, Aug 18 2026); OpenAI short-gaps preprint; arXiv 2608.03841; Adaptyv Bio and Twist Bioscience wet-lab validation; Lean 4 formalisations by Guo, Bajpai and Okur. Compiled by Dan's AI Intel.

A proof is not a proof until something that cannot be prompted agrees

Mathematics is the cleanest case, because mathematics has an unfakeable referee. A proof written in Lean either compiles or it does not, and the compiler has no view on whether the model that wrote it is impressive. That is why the Cycle Double Cover result is the strongest single item on the list: not because a model produced it in under an hour, but because three independent parties formalised it afterwards and the machine agreed with all of them.

It is also why the honest version of "dozens of conjectures cracked" is quieter than the headline. DeepMind's AlphaProof Nexus — a language model wrapped in an agentic loop around Lean — did settle nine open Erdős problems and prove 44 previously unproven conjectures from the Online Encyclopedia of Integer Sequences, for a few hundred dollars of compute. But it attempted 353 Erdős problems and 492 OEIS conjectures. The count is large because the pile is large.

Exhibit — Dozens of conjectures fell because thousands were attempted, not because most of them fell. A 2.5% hit rate on the hardest pile is still historic — but it is a search result, not a general solver. Source: Google DeepMind, AlphaProof Nexus preprint and formal-proof repository, 21 May 2026. Compiled by Dan's AI Intel.

The attribution is fragile in the other direction too. The Jacobian counterexample circulated widely on social media as a GPT-6 result and as evidence that OpenAI had reached AGI. It was Claude Fable 5, it was July, and the model that found it had been public for six weeks. Even the strongest results carry conditions the headline drops: Astra's prime-gap formalisation is verified in Lean, but conditional on three explicit input axioms, so "Astra solved prime gaps" is not a sentence the artefact supports. And the Erdős unit distance result disproved the bound Erdős conjectured; it did not resolve the underlying problem, which is still open.

The profession noticed. In June 2026, sixteen mathematicians published the Leiden Declaration on Artificial Intelligence and Mathematics, endorsed by the International Mathematical Union and signed on its first day by more than a thousand people including Terence Tao. It names five risks, and the ordering is instructive: unreliable results, missing citations, dependence on closed commercial systems, exaggerated claims, and loss of scientific independence. Only one of those is about the models being wrong. The rest are about a field losing the ability to check.

The same model scored 99.9% and 62.7% on the same test

Which brings us to why the launch tables are no longer usable, and this is worth sitting with because it is not a small technical caveat.

ARC-AGI-3 is the benchmark designed to resist memorisation: the model must find its way through an environment it has never seen. Astra scored 99.9% on it. It also scored 62.7% on it. The difference is not the model, the reasoning setting or the seed — it is the harness. Run through OpenAI's own provider adapter, which preserves hidden reasoning state between steps and compacts context automatically, the model finishes almost everything for roughly $19,000 of compute. Run through the ARC Prize foundation's provider-neutral rig, where the model has to carry forward only the notes it chose to keep, it manages 62.7% for roughly $26,000. The more expensive run scored thirty-seven points worse.

Exhibit — Thirty-seven points of the headline score came from the scaffolding, not the model. When the rig moves the number more than the model does, the number has stopped measuring the model. Source: ARC Prize Foundation, GPT-6 Astra results and accompanying analysis, September 2026. Compiled by Dan's AI Intel.

None of this is a scandal; OpenAI disclosed the setup, and the standardised 62.7% is still state of the art on its own — Astra used fewer actions than the median human tester on 96% of the levels it played, averaging 51.7% fewer. The ARC Prize foundation has itself said that saturating the benchmark would not constitute proof of AGI. But it makes a headline number uninterpretable without knowing the rig it was produced on, which is precisely the failure mode we walked through in the benchmark trap episode a few weeks back, and it is the same lesson as the harness episode: half of whether a model works was never the model.

The rest of the scoreboard is no steadier. On Humanity's Last Exam with tools, OpenAI's own launch table puts Astra at 57.2% and Fable 5.1 at 65.0% — the challenger publishing a table its rival wins. Artificial Analysis, measuring independently, ranks Fable 5.1 at 66 and Astra at 61, reversing the order of the science benchmark. And METR, whose task-completion time horizon has been the most respected single ruler for agentic capability, now attaches a warning to its own instrument: the strongest models it has assessed sit at roughly sixteen to twenty hours on the fifty-percent horizon, and measurements above sixteen hours are unreliable with its current task suite. The ruler ran out of ruler before the models ran out of capability.

Twelve weeks, and the score on doing science doubled

So if the scores are unreliable, what is the delta? It is real, and you can see it on two instruments that have nothing to do with each other.

The first is Terminal-Bench-Science 0.1 — seventy tasks contributed by working scientists across the life, physical, Earth, mathematical and engineering sciences, each requiring an agent to plan an investigation, run it in a terminal, and submit an artefact graded against hidden tests. It is not a quiz. Fable 5, released on 9 June 2026, scored 24.7%. Fable 5.1, released twelve weeks later, scored 52.6%. OpenAI reports 64.6% for Astra on the same benchmark. Every one of those figures is vendor-reported, with a standard error of three and a half to four and a half points, and the public leaderboard for the same benchmark tops out around 30% — so treat the levels with suspicion and the direction as sound.

Exhibit — The September models sit roughly twice as far up the science benchmark as the June fleet. The direction is the finding; the levels are vendor-reported and the public leaderboard reads far lower. Source: Anthropic launch table (Fable 5, Fable 5.1, Opus 5, GPT-5.6 Sol) and OpenAI launch table (GPT-6 Astra), Sep 1–3 2026. Terminal-Bench-Science 0.1, 70 tasks; vendor-reported standard error ±3.5–4.5 points. Compiled by Dan's AI Intel.

The second instrument is a freezer, and it agrees. In the August campaign the models' protein designs bound at 22.6% to 35.1%. With the September model the figure is close to 50% across twelve targets, with ten-fold affinity improvements on three of them. Nobody tunes a binding assay. Two independent measurements, one made of software and one made of protein, moved by roughly the same factor over roughly the same weeks. That is what a real capability jump looks like when you insist on checking it, and it is a far better signal than any single benchmark, because the two instruments have no way of flattering each other.

What changed is less mysterious than it sounds, and it is not raw reasoning. The distinctive thing about all six of the September results is that the model did not answer a question — it operated something. It drove a terminal, called open-source folding tools in a loop, parsed an instrument's undocumented binary format, trained a small neural network on decades-old radar returns. The step-change is in sustained, tool-using, self-checking work, which is exactly the axis METR's ruler stopped being able to measure.

Binding is not curing, and a map is not a landing

Now the honest half, because the gap between these two lists is the whole story.

Where the independent check is thin, the results are thin. The Venus map is a genuinely useful artefact, but nothing has yet ground-truthed it; that happens when VERITAS and EnVision arrive, years from now, and until then it is a very good inference from thirty-year-old radar, not a measurement. The kernel speedups are verified in the strongest sense available — bit-identical outputs — but the seven projects and the benchmarking are Anthropic's own. Anthropic itself states that its automated behavioural audit "provides less visibility into very long-context work and multi-agent settings," which is to say the least visibility exactly where the most capable work now happens.

And where there is no shortcut around biology, nothing has moved. As of July 2026, no drug whose target and molecule were both discovered by AI has been approved by the FDA. Not one. Molecules from AI-led discovery clear Phase I at 80 to 90% against a historical 52% — which mostly tells you they are chemically well-behaved. At Phase II, where the question stops being "does this bind" and starts being "does this help a person," the success rate is about 40% against a historical 37%: statistically indistinguishable. The translational gap — between predicting a molecule's properties and predicting its effect in a human body — is exactly as wide as it was. The furthest-advanced programme, Insilico Medicine's rentosertib for idiopathic pulmonary fibrosis, entered Phase III on 7 July 2026 and will not report for years.

There is a cheaper failure mode too, and it is the one to guard against personally: the claims travel faster than the corrections. A counterexample found by a three-month-old Anthropic model became, within a day, evidence on social media that OpenAI's newest model had reached AGI. The retraction never catches the original.

They are holding the best back, and now it says so on the box

We covered the suspicion that the labs sit on their strongest models in an episode last month. This generation converted that suspicion into published architecture, and the specific shape of it matters more than the fact.

Claude Mythos 5.1 is not a different model from Claude Fable 5.1. It is the same weights with the safeguards removed, released only to vetted cybersecurity teams and, through a Life Sciences Verification Program developed with the US government, to approved life-science organisations. OpenAI's life-sciences model, GPT-Rosalind, is available in one country, to vetted enterprises, through a Trusted Access Program whose launch partners are named: Amgen, Moderna, Novo Nordisk, the Allen Institute, Thermo Fisher Scientific and Dyno Therapeutics. And Astra is the first OpenAI model ever classified at the "Critical" cybersecurity tier of its own Preparedness Framework — it scored 100% on ExploitBench, discovered two previously unknown vulnerabilities during evaluation and chained them into a browser sandbox escape — so its cyber capabilities ship gated rather than open. Google did the same thing more quietly, shipping Gemini 3.8 Flash Cyber as a locked-down sibling on the same day as the general model. Three labs, one week, three gated twins.

Exhibit — The frontier did not move into a basement — it moved behind a vetting desk. The public benchmark and the actual frontier are no longer measuring the same artefact. Source: Anthropic, Claude Fable 5.1 and Mythos 5.1 launch note and Life Sciences Verification Program, 1 Sep 2026; Google, Gemini 3.8 Flash and 3.8 Flash Cyber, 2 Sep 2026; OpenAI, GPT-Rosalind trusted-access announcement and Preparedness Framework cyber assessment, Sep 2026. Compiled by Dan's AI Intel.

So the answer to "are they holding something back" is yes, but not in the shape the rumour has. There is no secret smarter model in a vault. There is the same model, twice, with a permission tier between the two copies — and the copy doing the drug design is the one the public will never score. What is genuinely unreleased is further up: Sam Altman says OpenAI expects an internal system that qualifies as AGI before the end of 2026, and chief scientist Jakub Pachocki says the company has already hit an internal benchmark for automating entry-level AI research work. Set against that, the most sobering line of the week came from OpenAI's own safety side, describing the monitoring it relies on to contain Astra's Critical-tier capabilities as "fragile" and "trending in a negative direction."

The five dials to re-read next check-in

The point of a recurring check-in is that it re-measures the same things, so here are the five dials, with today's reading. Not one of them is a leaderboard rank.

Exhibit — Five dials that a press release cannot move on their own. Four of the five require someone other than a model vendor to move them — which is the point. Source: Anthropic and OpenAI launch tables (Sep 2026); Adaptyv Bio and Twist Bioscience wet-lab results; Google DeepMind AlphaProof Nexus preprint (May 2026); METR Time Horizon 1.1; FDA approval records and published Phase II success rates as of July 2026. Compiled by Dan's AI Intel.

Read them together and the trajectory is legible without needing to settle whether Brockman is right. Two of the five moved sharply this generation, one moved in attempt count rather than hit rate, one broke as an instrument, and one has not moved at all in a decade of trying. That is not a smooth exponential and it is not a plateau. It is a capability that has become extremely good at the part of science that happens in front of a screen, and is still exactly as far as it ever was from the part that happens in a body.

Bottom line

The most important thing that happened in the first week of September 2026 was not that three labs shipped their best models, and it was not that one of them declared the AGI era. It was that the scoreboard everyone had been reading stopped working, and something better took its place: results that get checked by proof assistants, computer algebra, mass spectrometers and laboratory freezers — referees with no opinion about how the answer sounded.

Judge the next twelve weeks on that ledger. Ignore the launch tables, and ask instead what came back from outside the model's reach. Right now the answer is: rather a lot of mathematics, a surprising amount of protein, one map of Venus, and no medicine. And ask the second question too, because it will decide how much of this you ever get to see — whether the copy doing the most interesting work is the one you are allowed to run.

Sources

Transcript

Sam: Same model. Same test. Ninety-nine point nine percent.

Alex: And sixty-two point seven percent.

Sam: Nothing changed except the rig it ran on.

Alex: So if the scoreboard can move thirty-seven points on its own, what's left to measure?

Sam: Three hundred and fifty-four proteins that came back from a freezer in Switzerland having actually worked.

Alex: Welcome back to Dan's AI Intel — the show that takes the one question that actually matters and digs past the hype and the fear to what's really going on, and what it actually means. Not just for the technology, but for the money, the politics, and the race between the labs.

Sam: I'm Sam, and that's Alex, and today is a check-in. A recurring one, because this is a question we're going to keep coming back to: how fast are we actually accelerating toward superintelligence, and how would anyone outside these companies ever know?

Alex: The trigger is that in seventy-two hours at the start of September, every frontier lab on earth emptied the shelf. Anthropic, Google and OpenAI all refreshed their flagship models inside three days, and at one of those launches the president of OpenAI stood up and told the room, in as many words, that the AGI era has arrived.

Sam: Which is a genuinely enormous sentence to say out loud. And the first thing I wanted to do was go look at the benchmark tables, which is exactly what everybody did.

Alex: And that instinct is now broken. That's the deeper question underneath this whole episode — not is he right, but what would even count as evidence? Because the instruments everyone reads to answer that stopped working this generation, in a way I don't think most people have clocked yet.

Sam: So we're going to run the whole scan. Mathematics. Protein design. Analytical chemistry. A new map of Venus. The code that runs the models themselves. Every astonishing thing that came out of this generation — and then the much harder question of which of those things is actually true.

Alex: We'll show you the ladder we're using to sort them, we'll get into why the same model can post two scores thirty-seven points apart, and we'll end on five dials that we'll re-read every time we do one of these.

Sam: And there's one turn in here I did not see coming, about where the best models actually went. Not a rumour. It's printed on the box now.

Alex: Before we get into it — if you're new here, hit follow in whatever app you're listening in. It's free, and it means the next one just turns up.

Sam: Okay. Set the scene properly. What actually shipped, and when?

Alex: Between the first and the third of September, twenty twenty-six. On the first, Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1. On the second, Google released Gemini 3.8 Flash, and alongside it a locked-down security sibling called 3.8 Flash Cyber. On the third, OpenAI began rolling out GPT-6 Astra.

Sam: Three labs, three days. That's not a coincidence, that's a standoff.

Alex: And at the Astra briefing, OpenAI's president Greg Brockman said he personally believes the company has reached artificial general intelligence. And then he closed the briefing with five words. Welcome, he said, to the AGI era.

Sam: Right, and here's my honest reaction to that. I don't think he's lying. I think he might genuinely believe it. But belief isn't the thing I want. I want to know what would settle it.

Alex: That's exactly the right instinct, and the launch tables can't settle it. What I find much more interesting is the thing that happened in the same fortnight that got almost no coverage.

Sam: Which was?

Alex: For the first time, a meaningful volume of what these models produced got checked by things that are not models. A proof assistant. A computer algebra package. A mass spectrometer. A plate of engineered proteins sitting in a laboratory freezer.

Sam: Things with no opinion.

Alex: Things with no opinion. They're slow, they're unglamorous, and they are completely indifferent to how fluent the output sounded. And crucially, they're the only measuring stick left that a marketing department cannot tune.

Sam: So the move is to stop asking what got announced and start asking what survived being checked.

Alex: That's the whole episode. Measure the AGI curve the way you'd measure any other scientific claim.

Sam: Then let's do the scan, because I've been reading these headlines for a fortnight and half of them sound made up.

Alex: Most of them are real, and the raw list genuinely is startling. Start with mathematics. Five named open problems moved in four months.

Sam: Five. Named. Open problems.

Alex: In May, an OpenAI reasoning model disproved Erdős's unit distance conjecture, which dates from nineteen forty-six. Nine mathematicians, including the Fields medallist Timothy Gowers, then produced a human-verified version of the counterexample.

Sam: So humans went back over it afterwards and confirmed it holds.

Alex: In July, Claude Fable 5 produced an explicit counterexample to the Jacobian conjecture, in three variables, at degree seven. Levent Alpöge presented it, Terence Tao publicly digested it two days later, and the thing about a counterexample like that is you can check it in any computer algebra system in seconds.

Sam: That's the part I like. It's not a wall of prose you have to trust. It either works or it doesn't.

Alex: Then GPT-5.6 Sol Ultra, running sixty-four subagents in parallel, generated a proof of the Cycle Double Cover conjecture — open for fifty years — in under an hour. It's since been formalised in Lean 4 several times over by separate groups.

Sam: Sixty-four copies of itself working the problem at once. Quick aside — if the swarm thing grabs you, we did a whole episode on exactly that, AI Agent Swarm, number 54, just a few days ago.

Alex: And my favourite one. A neurosurgery resident in Beijing ran GPT-5.6 Sol for about sixteen hours and came out with a proof of Crouzeix's conjecture, open since two thousand and four. Michel Crouzeix checked it himself and thinks it's correct.

Sam: A neurosurgery resident. Not a maths department. That's a real shift in who gets to attempt these things.

Alex: And Astra improved a bound on short gaps between primes, from two hundred and forty down to one hundred and eighty-six, with a Lean formalisation attached.

Sam: Okay, that's mathematics. But maths still happens on a screen. What left the screen?

Alex: Protein design, completely. Across fifteen biological targets, Anthropic's models produced one thousand three hundred and twenty candidate binders. Adaptyv Bio and Twist Bioscience synthesised them and tested them in a lab, and three hundred and fifty-four of them bound, hitting fourteen of the fifteen targets.

Sam: Hold on. Give me the baseline. Is that good?

Alex: Hit rates ran from twenty-two point six percent to thirty-five point one percent depending on configuration. The industry norm is ten to fifteen percent.

Sam: So at the low end it beats the field, and at the high end it's more than triple it. On a first pass.

Alex: On one target, called RBX1, the model hit forty percent where the human entrants in a public design competition managed three point seven.

Sam: That's not an improvement, that's a different sport.

Alex: And with Mythos 5.1, the September figure is close to fifty percent across twelve targets — and on three of them the binding strengths came in ten times higher than the best designs any human submitted.

Sam: Right, and I want to be careful here, because binding is not the same as a drug. We'll come back to that hard.

Alex: We will, and you're right to flag it. Two more. Analytical chemistry: Claude was handed raw instrument files in undocumented vendor formats, worked out the encoding itself, and returned an NMR structure determination in twenty-three minutes and a purity analysis in nineteen.

Sam: Undocumented formats meaning nobody gave it the manual.

Alex: Nobody gave it the manual. And it landed within nought point nought eight of a hydrogen on the atom count, and called the purity at ninety-six point four percent against the contract laboratory's ninety-six point three three.

Sam: That's the sort of thing that quietly eats a job.

Alex: Then planetary science. Claude trained a neural network on radar data that NASA's Magellan orbiter collected more than thirty years ago, and produced a new elevation map of a third of Venus. Features resolved at two to three kilometres instead of ten to twenty, and height accuracy improved by up to twenty-five percent.

Sam: From data that has been sitting in an archive for three decades. Nobody collected anything new. The instrument was the model.

Alex: Anthropic released it under a Creative Commons licence, specifically so that NASA's VERITAS mission and Europe's EnVision can use it to choose what to photograph when they arrive.

Sam: That's a lovely detail, actually. It's not a demo, it's an input to somebody else's telescope time.

Alex: And last one, which is the strangest: computation itself. Mythos 5.1 rewrote the GPU kernels of seven open-source deep-learning models, hitting speedups of up to two and a half times, with bit-identical outputs — and cut the cost of genome-wide analyses by thirty to sixty percent.

Sam: Bit-identical meaning the answers didn't change at all, it just got there faster.

Alex: Exactly the same numbers out, two and a half times quicker. Anthropic says that's work which would normally occupy a team of performance engineers for weeks.

Sam: Okay, so six domains, and honestly all six sound roughly equally amazing when you say them in a row like that. Which is my problem.

Alex: And that's precisely the trap. They are not equally true, and the difference is not how impressive they sound.

Sam: So what's the sorting rule?

Alex: Four rungs. Announced by the lab. Method or artefact published. Re-checked by a machine or a referee. And then confirmed by something outside the model's reach entirely.

Sam: And every single one of these clears the first rung, because that's just a press release.

Alex: Right. Any of them can fill in the first column. Almost none of them can fill in the last one on their own. So read down the last column, not the first.

Sam: Let me try it. Mathematics and the protein binders go all the way to the end — the compiler agreed, the lab agreed.

Alex: They do, and notice what got them there: a Lean compiler in one case and a laboratory in the other. Neither of which the model can talk round.

Sam: The chemistry mostly does. The kernel speedups are strong but it's Anthropic checking Anthropic's own projects. The Venus map is a great inference nobody has been able to ground-truth. And drugs in humans is announced and published and then just stops.

Alex: That's the ledger. And once you see the results in that shape, the whole "are we near AGI" argument becomes a much calmer conversation.

Sam: Let's take the top rung seriously. Why is a proof assistant such a big deal here?

Alex: Because mathematics has a referee that cannot be charmed. A proof written in Lean either compiles or it doesn't, and the compiler has no view whatsoever on whether the model that wrote it is impressive.

Sam: It's a lock. The key turns or it doesn't. There's no rhetoric in the middle.

Alex: That's exactly it. And it's why the Cycle Double Cover result is the strongest single item on the entire list — not because a model produced it in under an hour, but because three separate groups formalised it afterwards and the machine agreed with all of them.

Sam: So the impressive bit isn't the speed. It's the audit.

Alex: The audit is the whole thing.

Sam: Alright, then let me put the headline I actually saw to you. Dozens of conjectures cracked. Is that true?

Alex: It's true, and it's quieter than it sounds. DeepMind's AlphaProof Nexus — which is a language model wrapped in an agentic loop around Lean — settled nine open Erdős problems and proved forty-four previously unproven conjectures from the Online Encyclopedia of Integer Sequences. For a few hundred dollars of compute.

Sam: Forty-four. That's the number that went round.

Alex: And it attempted three hundred and fifty-three Erdős problems, and four hundred and ninety-two OEIS conjectures.

Sam: Ah. So nine out of three hundred and fifty-three. That's about two and a half percent.

Alex: And forty-four of four hundred and ninety-two is under nine percent. The count is large because the pile is large.

Sam: I want to be fair to it though. A two and a half percent hit rate on the hardest open problems in the field, for a few hundred dollars, is still historic.

Alex: It absolutely is. It's just a search result rather than a general solver. Those are different claims about the world, and the headline collapsed them.

Sam: And this is the bit that annoys me, because the collapsed version is the one that travels.

Alex: It gets worse in the other direction too. The Jacobian counterexample went round social media as a GPT-6 result, and as evidence that OpenAI had reached AGI.

Sam: And it wasn't.

Alex: It was Claude Fable 5, it was July, and the model that found it had been public for six weeks.

Sam: So the wrong lab, the wrong month, and the wrong conclusion, on the most viral maths story of the year.

Alex: And even the correct results carry conditions the headline drops. Astra's prime-gap formalisation is verified in Lean, but conditional on three explicit input axioms — so "Astra solved prime gaps" isn't a sentence the artefact supports. And the Erdős result disproved the bound Erdős conjectured. It didn't resolve the underlying problem, which is still open.

Sam: Do the actual mathematicians have a view on all this? Because they're the ones being handed the work.

Alex: They wrote it down. In June, sixteen mathematicians published the Leiden Declaration on Artificial Intelligence and Mathematics. It was endorsed by the International Mathematical Union and signed on its first day by more than a thousand people, including Terence Tao.

Sam: What's in it?

Alex: Five risks, and the ordering is the interesting part. Unreliable results. Missing citations. Dependence on closed commercial systems. Exaggerated claims. And loss of scientific independence.

Sam: Only the first one is about the models being wrong.

Alex: Only the first one. The other four are about a field losing the ability to check.

Sam: That's a much more sophisticated worry than the one I expected, and it's the same worry we're going to hit in every other domain today.

Alex: Which brings us to the number we opened on, and I want to slow down here because it's not a small technical caveat.

Sam: This is the ninety-nine point nine and sixty-two point seven.

Alex: ARC-AGI-3 is the benchmark specifically designed to resist memorisation — the model has to find its way through an environment it's never seen. Astra scored ninety-nine point nine on it. It also scored sixty-two point seven on it.

Sam: Same model, same benchmark. So what moved?

Alex: Not the model, not the reasoning setting, not the seed. The harness. Run through OpenAI's own provider adapter, which preserves hidden reasoning state between steps and compacts the context automatically, the model finishes almost everything, for roughly nineteen thousand dollars of compute.

Sam: And the other run?

Alex: Through the ARC Prize foundation's provider-neutral rig, where the model has to carry forward only the notes it chose to keep, it manages sixty-two point seven — for roughly twenty-six thousand dollars.

Sam: Wait. The more expensive run scored worse?

Alex: Thirty-seven points worse.

Sam: Give me the picture, because that feels backwards.

Alex: Think of a chess player who's allowed to keep their own private notes on the board between every move, versus the same player who has to start each move from a blank sheet and only whatever they wrote down deliberately. Same player. Wildly different game.

Sam: And the notes aren't the intelligence.

Alex: The notes are the scaffolding around the intelligence. And when the scaffolding moves the number more than the model does, the number has stopped measuring the model.

Sam: I want to be fair here — is this a scandal? Did someone cheat?

Alex: No, and that matters. OpenAI disclosed the setup. And the standardised sixty-two point seven is state of the art on its own terms. Astra used fewer actions than the median human tester on ninety-six percent of the levels it played — averaging fifty-one point seven percent fewer.

Sam: So it's genuinely efficient, not just brute force.

Alex: And the ARC Prize foundation has itself said that saturating the benchmark wouldn't constitute proof of AGI anyway. The problem isn't dishonesty. It's that a headline number is now uninterpretable unless you know the rig it was produced on.

Sam: Which is exactly the thing we walked through in the benchmark trap episode, number 42, a few weeks back.

Alex: And it's the same lesson as the harness episode, number 48, from a couple of weeks ago — half of whether a model works was never the model.

Sam: Is the rest of the scoreboard any steadier? Please tell me something is stable.

Alex: It isn't. On Humanity's Last Exam with tools, OpenAI's own launch table puts Astra at fifty-seven point two percent and Fable 5.1 at sixty-five.

Sam: Sorry — OpenAI published a table where the competitor wins?

Alex: On that one, yes. Meanwhile Artificial Analysis, measuring independently, ranks Fable 5.1 at sixty-six and Astra at sixty-one, which reverses the order from the science benchmark.

Sam: So depending on which sheet you open, either lab is ahead.

Alex: And then the one that really got my attention. METR's task-completion time horizon has been the most respected single ruler for agentic capability — how long a task a model can carry to completion.

Sam: I know it. It's the one everyone plots as the doubling curve.

Alex: METR has now attached a warning to its own instrument. The strongest models it has assessed sit at roughly sixteen to twenty hours on the fifty percent horizon, and it says measurements above sixteen hours are unreliable with its current task suite.

Sam: So the ruler ran out of ruler before the models ran out of capability.

Alex: That's it exactly. It's a bathroom scale that stops at a number the thing you're weighing has gone past. The reading isn't wrong, it's just no longer a measurement.

Sam: And that's four independent ways the scoreboard failed in one fortnight. Harness dependence. Vendor tables that contradict each other. Independent rankings that reverse them. And the best ruler we had hitting its ceiling.

Alex: Which is why the whole exercise has to move somewhere else.

Sam: Okay, but I don't want us to overcorrect into nothing-is-knowable. If the scores are unreliable, is there actually a delta this generation, or not?

Alex: There is, and it's big. And you can see it on two instruments that have nothing to do with each other, which is the part that makes it credible.

Sam: Start with the first one.

Alex: Terminal-Bench-Science. Seventy tasks contributed by working scientists across the life, physical, Earth, mathematical and engineering sciences. Each one requires an agent to plan an investigation, run it in a terminal, and submit an artefact that gets graded against hidden tests.

Sam: So not a quiz. It has to actually do the work and produce something.

Alex: Fable 5, released on the ninth of June, scored twenty-four point seven percent. Fable 5.1, released twelve weeks later, scored fifty-two point six.

Sam: It doubled. In twelve weeks.

Alex: And OpenAI reports sixty-four point six for Astra on the same benchmark.

Sam: And now I have to ask the annoying question, because we've just spent ten minutes on this. Who measured that?

Alex: Good. Every one of those figures is vendor-reported, with a standard error of three and a half to four and a half points. And the public leaderboard for the same benchmark tops out around thirty percent.

Sam: Thirty. Against a claimed sixty-four.

Alex: So treat the levels with suspicion, and the direction as sound. That's the honest read.

Sam: You said two instruments. What's the second one?

Alex: The second instrument is a freezer, and it agrees.

Sam: The protein binders.

Alex: In the August campaign, the designs bound at twenty-two point six to thirty-five point one percent. With the September model it's close to fifty percent across twelve targets, with ten-fold improvements in binding strength on three of them.

Sam: And nobody tunes a binding assay. There's no harness argument to make. Either the protein stuck to the target or it didn't.

Alex: Two independent measurements — one made of software, one made of protein — moved by roughly the same factor over roughly the same weeks. That is what a real capability jump looks like when you insist on checking it.

Sam: Because the two instruments have no way of flattering each other.

Alex: None. And here's the bit I keep turning over: what changed is less mysterious than it sounds, and it isn't raw reasoning.

Sam: Go on, because everybody's default story is "it got smarter".

Alex: The distinctive thing about all six of those September results is that the model didn't answer a question. It operated something. It drove a terminal. It called open-source folding tools in a loop. It parsed an instrument's undocumented binary format. It trained a small neural network on decades-old radar returns.

Sam: So it's less "smarter" and more "it can now hold a long job together without falling over".

Alex: Sustained, tool-using, self-checking work. Which — and this is almost funny — is exactly the axis METR's ruler stopped being able to measure.

Sam: The thing that got good is the thing we lost the ability to score.

Alex: So now the honest half, because the gap between these two lists is the whole story.

Sam: Hit me. Where does it get thin?

Alex: The Venus map. Genuinely useful artefact, but nothing has ground-truthed it. That happens when VERITAS and EnVision arrive, which is years away. Until then it's a very good inference from thirty-year-old radar, not a measurement.

Sam: It's a forecast that hasn't met tomorrow yet.

Alex: The kernel speedups are verified in the strongest sense available — bit-identical outputs — but the seven projects and the benchmarking are Anthropic's own.

Sam: Marking your own homework, even if the homework is right.

Alex: And this one I think is the most quietly alarming line of the whole month. Anthropic itself states that its automated behavioural audit provides less visibility into very long-context work and multi-agent settings.

Sam: Which is where the most capable work now happens.

Alex: The least visibility exactly where the most capable work now happens. They said it themselves, in the launch material.

Sam: Right, and now I want the thing I flagged earlier. Medicine. Because this is the one everybody's family actually asks about.

Alex: Where there's no shortcut around biology, nothing has moved. As of July, no drug whose target and molecule were both discovered by AI has been approved by the FDA. Not one.

Sam: Not one. After how many years of this?

Alex: And the trial numbers are the thing to sit with. Molecules from AI-led discovery clear Phase I at eighty to ninety percent, against a historical fifty-two.

Sam: That sounds like a triumph.

Alex: Phase I is mostly safety. It mostly tells you the molecule is chemically well-behaved. At Phase II, where the question stops being does this bind and starts being does this help a person, the success rate is about forty percent — against a historical thirty-seven.

Sam: Forty versus thirty-seven. That's noise.

Alex: Statistically indistinguishable. The translational gap — between predicting a molecule's properties and predicting its effect in a human body — is exactly as wide as it ever was.

Sam: So go back to your lock. The key fits beautifully — and we still have no idea what's behind the door.

Alex: And opening the door is what Phase II is for. Which is why the honest answer here is a waiting one. The furthest-advanced programme is Insilico Medicine's rentosertib, for idiopathic pulmonary fibrosis. It entered Phase III on the seventh of July, and it won't report for years.

Sam: And meanwhile the headline is "AI designs drugs".

Alex: Which brings me to the cheapest failure mode of all, and it's the one to guard against personally. The claims travel faster than the corrections.

Sam: The Jacobian thing again.

Alex: A counterexample found by a three-month-old Anthropic model became, within a day, evidence on social media that OpenAI's newest model had reached AGI. The retraction never catches the original.

Sam: Okay, the turn you promised me. Where did the best models actually go?

Alex: We covered the suspicion that the labs sit on their strongest models about a month ago, in Frontier Labs Are Sitting On Their Best AI, number 39. This generation converted that suspicion into published architecture.

Sam: Published. As in they're telling us.

Alex: Claude Mythos 5.1 is not a different model from Claude Fable 5.1. It's the same weights, with the safeguards removed — released only to vetted cybersecurity teams and, through a Life Sciences Verification Program developed with the US government, to approved life-science organisations.

Sam: So it's not a smarter model. It's the same model with the handbrake off.

Alex: Same engine, two ignition keys. And OpenAI's life-sciences model, GPT-Rosalind, is available in one country, to vetted enterprises, through a Trusted Access Program whose launch partners are actually named: Amgen, Moderna, Novo Nordisk, the Allen Institute, Thermo Fisher Scientific and Dyno Therapeutics.

Sam: That's a very short list of the world.

Alex: And Astra is the first OpenAI model ever classified at the Critical cybersecurity tier of its own Preparedness Framework. It scored a hundred percent on ExploitBench, discovered two previously unknown vulnerabilities during evaluation, and chained them into a browser sandbox escape.

Sam: During the evaluation. Not in the wild, in the test.

Alex: So its cyber capabilities ship gated rather than open. And Google did the same thing more quietly — shipping Gemini 3.8 Flash Cyber as a locked-down sibling on the same day as the general model.

Sam: Three labs, one week, three gated twins. That's not three companies each having a policy idea. That's a shape.

Alex: And it answers the "are they holding something back" question, but not in the shape the rumour has.

Sam: Say the distinction, because I think this is the thing people will get wrong.

Alex: There is no secret smarter model in a vault. There is the same model, twice, with a permission tier between the two copies. And the copy doing the drug design is the one the public will never score.

Sam: Which means the benchmark everyone argues about and the actual frontier have quietly stopped being the same artefact.

Alex: They've stopped being the same artefact.

Sam: Is anything actually unreleased, then? Or is it all just permissions now?

Alex: What's genuinely unreleased is further up. Altman says OpenAI expects an internal system that qualifies as AGI before the end of this year. And the chief scientist, Jakub Pachocki, says the company has already hit an internal benchmark for automating entry-level AI research work.

Sam: Automating the job of the people who build it. That's the recursive one.

Alex: And set against all of that, the most sobering line of the week came from OpenAI's own safety side. They described the monitoring they rely on to contain Astra's Critical-tier capabilities as fragile, and trending in a negative direction.

Sam: Their words. About their own containment.

Alex: Their words.

Sam: I notice I have no way to check that one, either.

Alex: Which is the perfect place to land this, because the point of a recurring check-in is that it re-measures the same things. So here are five dials, with today's reading. Not one of them is a leaderboard rank.

Sam: Let's go. One.

Alex: The verified-science score. Today: fifty-two point six percent Anthropic-reported, sixty-four point six OpenAI-reported, and about thirty percent on the public leaderboard. What would move it: the independent leaderboard closing on the vendor tables.

Sam: Two.

Alex: The wet-lab hit rate on designed binders. Today: about fifty percent across twelve targets, against a ten to fifteen percent industry norm. What would move it: a second lab reproducing that on targets the model hasn't seen.

Sam: Three.

Alex: Machine-checked mathematics. Today: five named open problems in four months, and nine of three hundred and fifty-three Erdős problems attempted. What would move it: unconditional Lean proofs, and a rising hit rate rather than a rising attempt count.

Sam: That's the base-rate point made into an instrument. I like that.

Alex: Four. The independent agentic time horizon. Today: roughly sixteen to twenty hours at fifty percent reliability, and unreliable above sixteen. What would move it: a new task suite that can measure past a working day.

Sam: And five is the one that hasn't budged.

Alex: The distance from binding to curing. Today: zero FDA approvals for a fully AI-discovered drug, and Phase II at forty percent against thirty-seven historic. What would move it: one Phase II readout that beats the historical rate.

Sam: And here's what strikes me looking at all five together. Four of them require somebody other than a model vendor to move them.

Alex: Which is the point. That's the whole design.

Sam: So read them together, and what's the trajectory? Because I genuinely don't know how to summarise five dials that all did different things.

Alex: That's the honest answer, and you don't need to settle whether Brockman is right to see it. Two of the five moved sharply this generation. One moved in attempt count rather than hit rate. One broke as an instrument. And one hasn't moved at all in a decade of trying.

Sam: That's not a smooth exponential.

Alex: And it isn't a plateau either. It's a capability that has become extremely good at the part of science that happens in front of a screen, and is still exactly as far as it ever was from the part that happens in a body.

Sam: So let me try to say the whole thing in a sentence, because I think it changed how I'm going to read the next launch.

Alex: Please.

Sam: The most important thing that happened in the first week of September wasn't that three labs shipped their best models, and it wasn't that one of them declared the AGI era. It's that the scoreboard everyone had been reading stopped working — and something better took its place.

Alex: And the better thing is unglamorous. Proof assistants. Computer algebra. Mass spectrometers. Laboratory freezers. Referees with no opinion about how the answer sounded.

Sam: So the takeaway for the next twelve weeks is: ignore the launch tables.

Alex: Ignore the launch tables, and ask instead what came back from outside the model's reach. Right now the answer is: rather a lot of mathematics, a surprising amount of protein, one map of Venus, and no medicine.

Sam: And the second question, which I hadn't been asking at all before today.

Alex: Whether the copy doing the most interesting work is the one you're allowed to run. Because that's what decides how much of any of this you ever get to see.

Sam: That's the one that'll stay with me. And that's where we'll pick this up next time — same five dials, new readings.

Alex: And that's it for this check-in. Thank you so much for listening. I hope you come away with a sharper way of reading the next headline than you had an hour ago: it's a genuinely fast, genuinely complex picture with a brutally short knowability horizon, and that's exactly what makes it worth following closely rather than reacting to.

Sam: One honest note on how this show is made. It's AI-generated. Dan builds a custom stack of AI tools to research, analyse, verify and illustrate the questions worth understanding — mostly to learn them himself, and he publishes it for anyone who'd like to follow along. AI-assisted, fact-checked, and worth a second look.

Alex: Before we go, one genuinely useful thing you can do: follow the show. Whatever app you're listening in right now, there's a follow or a plus button. It's one tap, it's free, and it does two things. You'll get each new episode the moment it lands — and for a small independent show like this one, a follow is the single biggest lever there is for helping it reach other people trying to make sense of all this. So if this was worth your time, go ahead and hit follow.

Sam: And one last thing. This is a recurring check-in, which means what we measure next time is genuinely up for grabs. If there's a dial you think we've got wrong, or a result you want us to go and pull apart properly, tell us — podcast at connectiveshift dot com. We read every message, and it really does decide what we dig into next.

Alex: See you on the next one.