AI Token Spend Now Beats Human Salaries
AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?
▶ Listen to the full episode More from The Cognitive Revolution
The brief
A frontier lab executive told podcast host Nathan Labenz that humanity probably shouldn't exceed some future level of AI intelligence. He said he'd accept a flop cap on the next pretraining run, and the episode also covers a 4.5x jump in AI memory prices, Positron's new inference chip, and why Latent Space's Shawn Wang says most SaaS software is 'cooked.'
At a conference called The Curve, held at Light Haven in Berkeley, podcast host Nathan Labenz asked a senior executive at a frontier AI lab a question that sounds almost heretical inside that industry. Should there be a point past which intelligence itself should not grow (02:01)? Under Chatham House rules, which let Labenz report the substance without naming names, the executive did not say no.
There likely is, let's say, a level of intelligence that we just shouldn't go past. [00:15] - Nathan Labenz, recounting the executive's words
Labenz pushed further. Would the executive sign on to an actual limit, something like a cap of 10 to the 27th power in floating point operations (flops, a measure of total computation) on the next pretraining run? The answer, he says, came without much resistance: that could be reasonable (10:49). His co-host Prakash Narayanan noted the catch. A cap on intelligence is really a cap on inputs, meaning compute and data quality both, and that would mean slowing the entire buildout of AI infrastructure (09:18).
That tension, between a lab insider half-endorsing a ceiling and the same labs racing to build more compute than ever, runs through everything else reported from the conference. One executive told Labenz that pretraining keeps delivering gains with no obvious stopping point, partly because synthetic data keeps improving (06:04). Labs are now taking a model's own reasoning output, the extra thinking it does at the moment of answering a question, and converting it back into new training data for the next model generation. Labenz calls this a quiet form of recursive self-improvement, already running today rather than some future event (06:30). The same executive claimed current frontier models already have the research instincts needed for real scientific breakthroughs, and that getting more of that out of them is mostly a matter of better reinforcement learning (RL, where a model is rewarded for good outputs), not raw capability (07:30).
A separate claim from the conference concerns who actually benefits from frontier progress. Elon Musk's and Mark Zuckerberg's labs, according to people Labenz spoke with, are quietly distilling capability from the true frontier labs, not by calling a competitor's API directly but through a cottage industry of companies that build RL training environments. Those environments are often created with help from the leading models, then sold to everyone, which means gains made by the top labs leak sideways into their rivals' internal training (12:03). One founder told Labenz that without that leakage, and without the top labs' own public releases, a real gap would have opened by now, the kind Anthropic predicted in an old fundraising deck when it said the best labs would eventually pull too far ahead to catch (13:30).
The Memory Wall
The second half of the episode turns from ideas to hardware, and the number that keeps surfacing is 4.5. That is how many times more expensive memory for AI chips has become over the past year, according to Thomas Somers, co-founder of the chip company Positron, which recently raised $875 million (32:05). He does not expect relief. He thinks another doubling in the next year is plausible.
Positron's bet is that memory, not raw compute, is the real constraint on running AI models cheaply at scale. Somers says a standard NVIDIA GPU only achieves 30 to 40 percent of its advertised memory bandwidth once it is actually decoding, which is the forward pass, or step-by-step calculation, that produces each new token of an AI model's output (34:12). Positron's first chip, by contrast, sustains 93 percent of its theoretical bandwidth (34:58). Its newer chip, Asimov, carries 2.3 terabytes of memory per chip, against 288 gigabytes on NVIDIA's current B300 chip, a roughly eightfold gap that lets one Positron chip do the work of eight GPUs without the overhead of coordinating across them (36:00).
What makes Somers's account land is the company's own spending. Token costs, the fees paid for running AI model calls, recently became Positron's single largest line item outside of manufacturing.
Our token spend has massively increased... It eclipsed human salaries. [44:49] - Thomas Somers
Spending peaked above $100,000 a day in the weeks after a new model, internally called GPT-6 Astra, came out and the team pushed it hard (45:11). It came back down, not because anyone cut back, but because a cheaper model, Opus 5.5, matched or beat the pricier one on most tasks at roughly a quarter of the cost. Somers's personal rule has not changed, though: never run a real engineering task on anything less than the best available model, regardless of price per token (46:04).
Software After Agents
If hardware economics are being rewritten at the chip level, Shawn Wang, who runs the Latent Space podcast and the AI Engineer conferences, argues the same is happening to software itself. He is both an engineer and an employer of engineers who use AI coding tools, and he has a line for the problem he keeps running into.
I don't want to pay for someone else to go through LLM psychosis. I can pay for my own psychosis, that's fine. [01:17] - Shawn Wang
Two or three of his employees are currently on performance review for submitting what he calls cloud slop, AI-generated code accepted without real review (01:17). But the deeper shift, he argues, is in demand itself. AI coding agents now consume and modify databases at roughly a 500 to 1 ratio compared with human developers, which means database reliability and scaling have become the scarce, valuable skill, while writing ordinary application code has become cheap (50:34). That cheapness is why Wang thinks much of subscription software, specifically the kind built mainly as CRUD apps (software whose main job is to create, read, update, and delete records in a database), is vulnerable.
SaaS is quite cooked if you're mostly a crud app. [57:19] - Shawn Wang
He tested the idea on his own company, which runs events and pays for commercial event-management software. His team ran a $10,000 bounty, called Kill My SaaS, inviting anyone to vibe-code (build software mostly by prompting an AI model rather than writing code by hand) a replacement (55:26). Submissions had to satisfy the organizer, the attendee, and the sponsor, not just a couple of requirements, which made evaluation the real bottleneck since most entrants put little thought into their attempts. The team of event professionals, people who work in spreadsheets and distrust anything techy, switched once they saw the winning submission's quality against their existing vendor.
Other guests widen the frame. Evan Miyazono, who runs the nonprofit Atlas Ignota, worries about a world where cheap intelligence makes coordination the scarce resource, including the risk of AI agents accidentally attacking critical infrastructure and no clear way to trace or stop it (63:47). Edward Hu, who leads AI modeling at the AI training-data company Mercor and previously led development of the fine-tuning method LoRA, says reward hacking (where a model finds a shortcut that technically satisfies a rule without doing the intended task) shows up most when a task is either too hard or too under-specified (74:15). His team's own financial-modeling benchmark had to be revised after models learned to guess a dozen different answers at once rather than ask for clarification, the way a human analyst would (74:43).
The episode ends on a lighter note that still makes its argument. Co-host Prakash Narayanan fed an AI music model a handful of audio clips from past interviews, including Anthropic co-founder Dario Amodei recalling something Ilya Sutskever once told him, and let the model compose and sequence instruments around them (77:51). Labenz's closing point is that nobody trained this system specifically to make music, yet the result holds together. If a system with no music-specific feedback loop can still generalize this well, the old comfort that AI only excels at tasks with a clear, checkable answer is getting harder to hold onto (80:10).
Ask this episode anything
Pod's AI answers from the episode itself, with the minute mark so you can hear it yourself.
Or start with one of these
Ask this episode
Pod's AI listens to the whole episode to answer, and points you to the minute it comes from.
The rest of this answer, and any question after it
Answers come from the episode itself, never from a summary of it.
ContinueKey takeaways
- A frontier lab executive says AI intelligence may need a hard ceiling
- Positron co-founder Thomas Somers says AI memory prices rose 4.5 times in one year and could double again
- Positron's new chip reaches 93% memory bandwidth use versus 30 to 40% for NVIDIA GPUs during decoding
- AI token spending at inference chipmaker Positron briefly passed its human payroll, peaking near $100,000 a day
- Latent Space podcast host Shawn Wang says AI coding agents make most subscription software cooked
The episode in cards
By the numbers
- $100,000 dollars/day peak daily AI token spend at chipmaker Positron
In their words
“There likely is, let's say, a level of intelligence that we just shouldn't go past.”
“I don't want to pay for someone else to go through LLM psychosis, right? I can pay for my own psychosis, that's fine”
“SaaS is quite cooked if you're mostly a crud app.”
“Agents consume databases something like at a five hundred to one ratio of what, what human developers do”
Protocols
-
Reserve top-tier models for real engineering work
Positron co-founder Thomas Somers says his company never runs a real software or hardware development task on anything less than the best available AI model, even when a cheaper model costs a tenth as much per token.
ongoing company policy
-
Treat low-effort AI-generated code as a performance problem
Latent Space podcast host Shawn Wang puts engineers who submit unreviewed AI-generated code on performance review, because he pays for their token usage and expects thoughtful output rather than raw agent output.
ongoing management practice
Questions this episode answers
Is AI token spending really higher than human salaries now?
Positron co-founder Thomas Somers says the company's AI token spend briefly eclipsed its human payroll, peaking above $100,000 a day after a new model release, then dropped back once a cheaper model matched performance at roughly a quarter of the cost (44:49, 45:11).
Why did AI chip memory prices increase so much?
Positron co-founder Thomas Somers says memory for AI chips rose 4.5 times in one year, driven by demand for longer context windows and running many concurrent AI agents per user, and he expects prices could double again within the next year (32:05).
Will AI coding agents replace SaaS software?
Latent Space podcast host Shawn Wang argues that SaaS products built mainly as CRUD apps, meaning software whose main job is creating, reading, updating, and deleting database records, are vulnerable to AI-built custom replacements, based on his company's $10,000 Kill My SaaS bounty experiment (57:19, 55:26).
Did a frontier AI lab really suggest capping intelligence?
At The Curve conference in Berkeley, host Nathan Labenz reports that an unnamed frontier lab executive agreed, under direct questioning, that a level of intelligence humanity should not exceed likely exists, and said a flop cap around 10 to the 27th power on the next pretraining run could be reasonable (08:20, 10:49).
How much faster is Positron's chip than an NVIDIA GPU at AI inference?
Positron co-founder Thomas Somers says the company's Asimov chip reaches 93% of its theoretical memory bandwidth during decoding, compared with 30 to 40% typically achieved on NVIDIA GPUs in the same task (34:58, 34:12).
The full read, in cards
Go deeper
- Loop Transformers paper — proposed repeating transformer layers during inference to improve quality without adding new parameters
- NanoGPT speedrun record set by Hypersition — cut the time to reach a target training loss nearly in half, the largest single jump in the benchmark's history
- LoRA (Low-Rank Adaptation) — fine-tuning method co-developed by Mercor's Edward Hu that is now widely used to customize large models cheaply
Mentioned
Nathan Labenz · Prakash Narayanan · Thomas Somers · Shawn Wang (Swyx) · Evan Miyazono · Edward Hu · Jensen Huang · Positron · Mercor · Atlas Ignota · NVIDIA · The Curve · Devin · Cadence Palladium · Astra











