The Cognitive Revolution artwork

The Cognitive Revolution

Why Hyperscaler GPU Prices Beat Neo-Clouds

AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases

▶ Listen to the full episode More from The Cognitive Revolution

The brief

This roundup links three threads: US-China AI verification hinges on unfakeable signals like data center heat, hyperscalers charge 2-3x neo-clouds for the same GPUs because of bundled enterprise stickiness, and agentic AI now scores 50% on a rare-disease genome benchmark versus 10% for older tools.

A data center throws off heat that nobody can hide. That single physical fact runs under most of this week's conversations. Jeremy and Ed Harris, researchers at the AI policy group Gladstone AI, have spent months interviewing roughly a dozen State Department diplomats who negotiated with China. Their question: if an AI incident happens tomorrow, what could the United States credibly ask China to stop doing, and what could anyone actually verify? Jeremy Harris's answer starts with thermodynamics. "Data centers put off a hell of an energy footprint," he says. "The thermals on those are really, really bright" (06:12). A cluster above a certain size cannot be disguised, at least not yet, which makes its heat signature a kind of involuntary honesty.

The problem is that this honesty is also a blunt instrument. The brothers walk through the logic: trust between Washington and Beijing is thin, AI capability keeps advancing, and when something alarming finally happens, the only lever available by default will be an extreme one, telling the other side to shut down every large cluster (07:21, 08:26). That is an expensive, almost unusable threat in either direction, since GPU depreciation is already the dominant cost on a data center's books. Better verification technology would narrow the ask. If monitoring tools improve by even half, Jeremy Harris argues, a crisis response might only require shutting down half a data center instead of all of it, "saving billions of dollars right off the bat" (09:25). The catch is that new verification methods typically take years to be vetted by the intelligence community into what's called a national technical means, the formal category of tool allowed to underpin a treaty-like judgment (11:09). Harris wants the small community of verification startups and their intelligence-community counterparts to already have each other's phone numbers before a crisis arrives, so that the bureaucratic lag does not become the bottleneck.

Reading diplomatic language is its own kind of signal detection. Xi's speech at the World AI Conference struck the Harrises as conciliatory: he credited the US with inventing AI and warned against overstretching the concept of national security (20:21). They are careful not to over-read a speech, since rhetoric is often strategic theater. More convincing would be something harder to fake, like China's AI regulator publicly disciplining a model for ideological drift the way US labs discuss their own safety incidents (22:09). The underlying idea, one Ed Harris keeps returning to, is that certain kinds of transparency cost little and buy real trust.

"The kind of transparency that goes like, 'Hey, we're giving you enough vision into what we're doing to see that we are not doing the thing you fear most,' that sort of thing is potentially useful." — Ed Harris, Gladstone AI [00:24]

Signals also show up in markets, just priced in dollars instead of heat. Steve Ho, head of research at SiliconData, which builds price indexes for rented GPUs and model tokens, was asked why an identical chip costs two to three times more at a hyperscaler like AWS or Azure than at a smaller neo-cloud. It isn't inefficiency. It's bundling. Hyperscalers wrap the raw chip in software, compliance tooling, and years of enterprise relationship, and customers pay a premium for not having to switch.

"They charge regularly, consistently, at least two to three times, sometimes more, compared to a typical neo cloud… It is being sold as very much a differentiated product." — Steve Ho, SiliconData [35:53]

The price of being first

Ho's team also tracks a token expenditure index that recently ticked up 11.5% over seven days (38:48), which sounds like AI is getting pricier at exactly the moment everyone insists it is getting cheaper. Both things are true. Per-token prices on frontier models have fallen roughly in half since late June, from about $4 to $1.60 per million tokens (40:24). But the index is expenditure-weighted, so when users shift their spending toward a pricier, more capable model, the index can rise even as sticker prices keep falling (41:12). It is measuring a change in taste, not a change in cost.

Speed is becoming its own selling point. At OpenAI's developer event, the new Ultra Fast mode runs eight times quicker than the standard release, fast enough to build software interactively as a user types (44:30). A separate feature, letting a user sign in to any partner app with their existing ChatGPT account, means a startup no longer has to burn tokens profiling every new signup before it can show value, cutting a real cost of customer acquisition (45:56). Novelist Joel Borgen, who co-wrote his book "The Receipt Horizon" with AI models, offers the more sobering half of this story: raw model prose, even with a detailed outline, still comes out unreadable without a human who knows how to scaffold and edit it (02:17).

What a genome will not say for itself

The most consequential signal in this episode is biological. Daniel McKinnon's son Owen died of a rare lung disease. A standard whole-genome sequencing lab had analyzed Owen's genome and found nothing. A human specialist later found the cause: a 91-kilobase deletion, a missing stretch of DNA, that knocked out an enhancer, a region that switches a gene on, located a million bases away from the gene itself (61:12). The lab's own filtering rules, reasonable under time pressure, excluded structural variants more than one kilobase from the gene in question. McKinnon later built his own AI pipeline, looping an early agentic model (OpenAI's o3) repeatedly over the same data until it, too, surfaced the deletion his family's original lab had missed (62:17).

That experience became Gamo Labs, a company that uses AI agents to reanalyze genomes stuck in diagnostic limbo. McKinnon says most whole genomes done for infants suspected of a genetic disorder come back non-diagnostic, not for lack of information but for lack of time: a human lab has to draw a line somewhere and move on (62:55).

"Most whole genomes, even with infants suspected to have genetic disorders, come back non-diagnostic. And I really think it's one of these like meat space problems." — Daniel McKinnon, Gamo Labs [62:55]

Genetic variants that can't yet be classified sit in a holding category called a variant of uncertain significance, scored against a rubric from the American College of Medical Geneticists, or ACMG. A variant needs six points to be reclassified as likely pathogenic, which unlocks treatment and insurance options (65:09). Gamo Labs built a benchmark called Rarebench to measure how well different tools rank candidate variants. An established machine-learning tool called Lyrical scores about 10% on it; Anthropic's Claude Code, used with no special tuning, scores about 50% (69:53). McKinnon's company adds a harness around the model, custom tool routing, and close trace analysis to catch the kind of error a frontier model can still make, like silently renaming a gene mid-analysis (74:27). The team recently took over a biology lab to run functional experiments on uncertain variants directly, closing the loop between a model's prediction and a real cell line (66:33, 79:23).

McKinnon is candid that this progress has a cost on the other side of the ledger too. If frontier labs slow down to manage AI risk, as the earlier conversation about verification suggests they might, some fraction of unresolved genetic cases will stay unresolved longer. The 50% score on Rarebench is itself a reminder: half the questions families bring to his company still don't get an answer (82:48).

"We are diagnosing kids today. We have many examples of kids who are undiagnosed who we have diagnosed at a small company four months in." — Daniel McKinnon, Gamo Labs [81:27]

None of the week's threads resolve cleanly. The Harrises want verification technology that doesn't yet exist, built fast enough to matter. SiliconData is pricing a market that is still inventing its own units. McKinnon is racing a benchmark score toward 100% knowing that each missing point is a family waiting on an answer. What connects them is a simple discipline: look for the signal that would be expensive or impossible to fake, whether it's heat from a data center, a price that moves with usage mix, or a deletion a filter was never built to see.

Ask this episode anything

Pod's AI answers from the episode itself, with the minute mark so you can hear it yourself.

Or start with one of these

Key takeaways

  • Hyperscalers charge 2 to 3 times more than neo-clouds for the same GPU, mostly for bundled enterprise stickiness, not better chips
  • New AI verification technology for US-China trust takes years to become usable national technical means
  • A token price index can rise 11.5% in a week even as per-token prices fall, because usage shifts toward pricier models
  • Claude Code scores about 50% on Gamo Labs' Rarebench genome benchmark versus 10% for older machine learning tools
  • A genetic variant needs six ACMG rubric points to move from uncertain to likely pathogenic, unlocking treatment options
How Gamo Labs reopens a missed genetic diagnosis — "The Cognitive Revolution" : AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases

The episode in cards

Hyperscaler cloud versus neo-cloud GPU rental — "The Cognitive Revolution" : AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases

By the numbers

  • 6 points ACMG rubric score needed to reclassify a gene variant as likely pathogenic [65:09]
  • 50% percent score Claude Code achieves on the Rarebench rare-disease genome variant benchmark [69:53]

In their words

“Certain kinds of transparency can be stabilizing. So the, the kind of transparency that goes like, "Hey, we're giving you enough vision into what we're doing to see that we are not doing the thing you fear most."”

Ed Harris [00:24]

“Data centers put off a hell of an energy footprint. You know, the thermals on those are really, really bright. Data centers are huge.”

Ed Harris [06:12]

“Uh, they charge regularly, consistently, at least two to three times, sometimes more, compared to a typical, a neo cloud”

Steve Ho [35:53]

“Most whole genomes, even with infants suspected to have genetic disorders, come back non-diagnostic. And I really think it's one of these like meat space problems.”

Daniel McKinnon [62:55]

“Years." We're diagnosing kids today. We have many examples of kids who are undiagnosed who we have diagnosed at a small company four months in”

Daniel McKinnon [81:27]

Protocols

  1. Request raw genome data after a non-diagnostic result [60:03]

    Daniel McKinnon, founder of Gamo Labs, says families whose whole genome sequencing comes back non-diagnostic can call the sequencing lab and request their raw data under HIPAA rights, which then allows an independent reanalysis by another lab or AI pipeline.

    once, after a non-diagnostic whole genome result

  2. Add sign-in with ChatGPT to cut onboarding token costs [45:56]

    Podcast co-host Prakash recommends that SaaS companies implement sign-in with ChatGPT so new users bring their existing plan, which removes the need to spend tokens profiling every free-trial signup before the product shows its value.

    at signup, for new trial users

Questions this episode answers

Why are hyperscaler GPU prices higher than neo-cloud GPU prices?

SiliconData's head of research Steve Ho found hyperscalers charge 2 to 3 times more than neo-clouds for the same chip because the price bundles enterprise software, compliance tooling, and long-standing customer relationships, not because the hardware performs differently (35:53). The premium reflects product differentiation and switching stickiness rather than inefficiency.

How can the US verify China is not crossing AI red lines?

Gladstone AI researchers Jeremy and Ed Harris describe data center heat signatures as a near-term verification method, since large GPU clusters emit energy footprints visible with existing national technical means (06:12). They caution that new verification technologies typically take years to be vetted by the intelligence community before they can be relied on in a crisis (11:09).

How well does AI perform at diagnosing rare genetic diseases?

On Gamo Labs' Rarebench benchmark for prioritizing genome variants, Claude Code scores about 50% compared to roughly 10% for an established machine learning tool called Lyrical (69:53). Founder Daniel McKinnon notes this still leaves about half of cases unresolved, since most whole genomes for infants with suspected genetic disorders return non-diagnostic from standard labs (62:55).

What is a variant of uncertain significance in genetic testing?

It is a genetic finding that cannot yet be classified as harmful or harmless. Daniel McKinnon explains that labs score these variants against the American College of Medical Geneticists rubric, and a variant needs six points or more to be reclassified as likely pathogenic, which then unlocks treatment and insurance options (65:09).

Why did a token price index rise while per-token AI prices were falling?

SiliconData's proprietary LLM expenditure index rose 11.5% over seven days even as headline per-token prices fell by more than half since June, because the index is expenditure-weighted (38:48, 40:24). Steve Ho explains that when users shift spending toward more expensive, powerful models, the index rises even without any price increase (41:12).

The full read, in cards

Go deeper

  • Gladstone AI diplomat report — a newsletter summarizing interviews with about a dozen State Department diplomats on verification and red lines with China [04:35]
  • AI 2027 — forecast referenced for how quickly data center concealment timelines could shorten [06:12]
  • The Receipt Horizon — Joel Borgen's novel, set after an AI war, written in collaboration with AI models [47:40]
  • Rarebench — Gamo Labs' benchmark showing Claude Code scoring about 50% versus 10% for the older tool Lyrical [69:53]

Mentioned

Gladstone AI · Jeremy Harris · Ed Harris · SiliconData · Steve Ho · Gamo Labs · Daniel McKinnon · Joel Borgen · Rarebench · ACMG · The Receipt Horizon · OpenAI · Anthropic · Claude Code · Counsel