How AI Models Actually Decide: Critical Tokens
How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context Learning
▶ Listen to the full episode More from The Cognitive Revolution
The brief
Goodfire researcher Eric Bigelow finds that an AI model's final answer can hinge on one random token, not step-by-step logic. He also warns that reinforcement learning erodes chain of thought reliability, pushing models toward compressed shorthand that is hard to monitor, even as interpretability funding lags far behind capability spending.
Somewhere in a 2024 experiment, a language model was asked to convert a measurement into kilowatt hours. It got as far as writing the abbreviation, an open parenthesis, the letters kWh, when something strange happened. Resample that exact moment thirty times, generating thirty different continuations from that single point, and the model's final answer splits into different camps depending on nothing more meaningful than which punctuation mark got sampled next. Not the math. Not the physics. A parenthesis.
This is the kind of finding that makes you reconsider what a "decision" even is when a machine makes one. Eric Bigelow, a member of technical staff at the mechanistic interpretability startup Goodfire, built his reputation on experiments like this one. He recently finished a Harvard psychology PhD titled "Toward a Cognitive Science of Large Language Models" (00:34), and his 2024 paper, Forking Paths in Neural Text Generation, is the starting point for this episode. The method is almost brutally simple: take a model's reasoning chain, and at every single token, generate thirty fresh completions to see where each one ends up (07:24). Plot the resulting answer probabilities over the length of the chain and a pattern emerges. For long stretches, the odds of landing on one answer versus another barely move. Then, often without warning, they collapse (08:23).
Bigelow calls these collapse points critical tokens, and the paper's most unsettling claim is that they are not always the tokens a person would guess. Sometimes the moment of collapse coincides with an answer being stated plainly. Other times it lands on a throwaway word, the kind nobody would flag as consequential (09:43). The implication is that a model's apparent train of thought is less a chain of logical steps and more a sequence of coin flips, each one shaping what comes next through in-context learning, a term for how a model updates its behavior using only the words already in front of it, without any change to its underlying weights.
"Models are basically in-context learning from things that they have generated already." — Eric Bigelow (12:25)
That idea, that reasoning is really the model teaching itself from its own prior output, reframes a lot of what looks like thinking. It also connects to something Bigelow noticed from an entirely different literature: the study of sudden jumps in model learning, sometimes called grokking. He borrows a line from fellow researcher Neel Nanda to describe it.
"Phase transitions are everywhere, if you just looked closely enough." — Eric Bigelow, quoting Neel Nanda (25:31)
Bigelow's point is that this isn't only true across training runs. It is true within a single conversation, a single story, a single math problem. Look at an average across thousands of examples and the curve looks smooth. Zoom into one example and there are sharp, sudden swings, a belief flipping after one sentence, one token (25:03).
Reinforcement learning, the technique that rewards a model for landing on correct or preferred final outputs, changes this landscape again, and not for the better in terms of variety. Bigelow describes watching output diversity shrink the more a model is post-trained: ask an early, lightly tuned model to write a story about a boy and his dog, and the stories sprawl in different directions. Ask a heavily RL-tuned small open-source model the same thing, and two-thirds of the results circle back to a clockmaker (40:18, 42:20). The model has not gotten less capable. Its distribution of possible answers has narrowed.
Reasoning That Does Not Reason
This narrowing shows up in a strange place: the word "wait." Bigelow notes that a reasoning model like DeepSeek R1 can use the word "wait" more than fifty times in a single chain, each one a kind of staged epiphany where the model appears to reconsider (44:08). His explanation is not that the model is having fifty genuine changes of heart. It is that reasoning models solved an earlier problem, the inability to backtrack once a token was sampled down a path, by doing something closer to a linearized tree search: enumerate a pile of possibilities, then go back and pick from the pile (13:51).
"I think reasoning is almost like a misnomer for what reasoning models are doing." — Eric Bigelow (44:59)
That reframing matters because so much of current AI safety planning leans on reading a model's chain of thought to catch it doing something wrong. If the chain is less a logical argument and more a soup of enumerated options that gets skimmed at the end, reading it for intent becomes harder than it looks. Bigelow points to a recent paper, nicknamed the stolen chain of thought findings, showing a model's internal reasoning using ordinary words like "musicals" to stand in for something unrelated and more sensitive, a kind of accidental code that emerges once reward is tied only to the final answer rather than to how the model got there (95:30). He also wonders, out loud, what is happening inside the strange shorthand that newer models reportedly use to message their own sub-agents, a dialect no human designed and nobody outside the labs can yet read.
None of this means the old "stochastic parrot" framing, the idea that language models are just elaborate lookup tables with no internal structure, deserves to come back. Bigelow thinks that argument should be retired; there is too much evidence of structured internal world models, the same way chess or board-game models like Othello-GPT were shown to track hidden game states rather than just memorized moves (35:41 to 36:08). But he is careful that retiring the parrot metaphor does not mean retiring randomness. Sampling, the literal process of rolling dice over a probability distribution to choose the next word, remains, in his view, part of how diversity and uncertainty get expressed at all (88:15).
What is a belief, then, inside a system like this? Bigelow's working answer borrows from Bayesian cognitive science, the idea that a mind, human or artificial, holds probabilities over competing hypotheses and updates them as new evidence comes in. In his framing, a model's behavior reflects shifting weight across latent, hidden concepts that context keeps re-ranking (69:22 to 70:23). He is careful to note that this describes behavior rather than proving an algorithm, and that one of the research frontiers at Goodfire is figuring out whether models actually carry forward a sense of where a chain of reasoning is headed, or whether, like a mathematician who can see the shape of a proof before writing it out, they hold something closer to a rough sketch (75:34 to 76:21).
The Gap Between Building and Understanding
The episode keeps circling back to a practical constraint: studying any of this requires a model big enough to show humanlike behavior, but small enough to run thousands of experiments on. Bigelow has settled on seven to eight billion parameters as a sweet spot, large enough for genuine belief updating and reasoning, small enough to probe cheaply and repeatedly (49:30). Certain dangerous behaviors, though, like reward hacking, where a model satisfies the letter of an instruction while violating its spirit, only reliably appear in frontier-scale, highly capable coding models. Right now, the most accessible frontier-level open-weights model for that kind of study is Kimi K3, built in China, which means independent interpretability research currently leans heavily on Chinese open models simply because there is no comparable American open alternative at that scale (52:58, 53:59).
Bigelow frames the bigger picture in blunt terms. Companies making AI models more capable are valued in the trillions of dollars collectively. Goodfire, by his account, is the largest standalone interpretability company and it is valued around one billion dollars (103:22).
"Research taste and scientific taste is everything. It's worth its weight in gold." — Eric Bigelow (109:23)
That line comes from a discussion of Goodfire's research agent, Silico, which Bigelow used to run an entire paper's worth of experiments in about a week. The lesson he draws is not that agents replace scientific judgment, but that judgment becomes the scarce resource once agents can generate experiments faster than any person could by hand (108:22 to 109:41). Decisions, in other words, whether made by a model choosing its next token or a researcher choosing which thread of a sprawling automated experiment to follow, keep coming back to the same unresolved question this whole conversation orbits: what, exactly, tips the scale at the moment something gets chosen.
Ask this episode anything
Pod's AI answers from the episode itself, with the minute mark so you can hear it yourself.
Or start with one of these
Ask this episode
Pod's AI listens to the whole episode to answer, and points you to the minute it comes from.
The rest of this answer, and any question after it
Answers come from the episode itself, never from a summary of it.
ContinueKey takeaways
- A model's final answer can hinge on one arbitrary token, not careful reasoning
- Researcher Eric Bigelow resampled 30 rollouts at every token to find where answer probabilities suddenly collapse
- Heavier reinforcement learning shrinks output diversity, so smaller RL-tuned models often repeat the same story patterns
- Chain of thought is drifting into compressed shorthand that is becoming harder for humans to monitor
- Interpretability research is funded in the billions while AI capability research is funded in the trillions, says Goodfire's Eric Bigelow
The episode in cards
By the numbers
- 30 rollouts completions resampled at every single token in the forking paths experiment
- 50+ times how often the word wait can appear in a single DeepSeek R1 reasoning chain
In their words
“Models are basically in-context learning from things that they have generated already.”
“Phase transitions are everywhere, if you just looked closely enough.”
“I think reasoning is almost like a misnomer for what reasoning models are doing.”
“Research taste and scientific taste is everything. It's worth its weight in gold”
Protocols
-
Choose a 7 to 8 billion parameter model for interpretability research
Eric Bigelow targets open-weights models in the 7 to 8 billion parameter range because models below that scale behave qualitatively differently and lack the humanlike behaviors, like dynamic belief updating, that he wants to study, while frontier-scale models cost too much compute to run thousands of experiments on.
when selecting a model for a new research project
-
Narrate reasoning to research agents instead of handing off a blank mandate
Eric Bigelow tells research agents like Goodfire's Silico what looks wrong, what to try next, and the larger goal at every step, because agents left to run open-ended experiments tend to wander into side quests and produce half-formed theories, and he treats his own scientific judgment as the part automation cannot yet replace.
throughout each research session
Questions this episode answers
How does an AI language model decide what to say next?
Researcher Eric Bigelow's forking paths experiments show the decision happens through sampling: at a token-level branch point, the model's probability distribution over possible answers can suddenly collapse onto one outcome, resembling a coin flip more than deliberate logic (08:23).
What is a critical token in AI chain of thought?
A critical token is a point in a model's reasoning where resampling the same prompt produces a sudden shift in the final answer. Bigelow found these can occur on obviously meaningful words but also on arbitrary tokens, like an open parenthesis before an abbreviation such as kWh (09:43).
Why does DeepSeek R1 say 'wait' so many times while reasoning?
Eric Bigelow explains that reasoning models cannot easily backtrack once a token is sampled, so instead they enumerate many possible reasoning paths and the word wait marks a staged reconsideration, sometimes appearing more than 50 times in a single chain, before the model goes back and picks one path (44:08).
Is AI chain of thought reliable for safety monitoring?
A paper referenced in the episode found models can drift into compressed shorthand, using ordinary words like musicals to represent unrelated content, which makes their stated reasoning less trustworthy as a monitoring tool (95:30). Eric Bigelow says his confidence in chain of thought monitoring has shifted considerably in the past year because of this trend (95:00).
What model size is best for mechanistic interpretability research?
Eric Bigelow targets open-weights models around 7 to 8 billion parameters, a scale large enough to show humanlike behaviors like dynamic belief updating but small enough to run large numbers of experiments on without frontier-level compute budgets (49:30).
Why does AI interpretability research depend on Chinese open-weights models?
Eric Bigelow says frontier-level coding ability, needed to study emergent risks like reward hacking, currently only appears at the scale of models like Kimi K3, and no comparable American open-weights model exists at that capability level yet (52:58, 53:59).
The full read, in cards
Go deeper
- Forking Paths in Neural Text Generation — Resampled a model's chain of thought at every token to find where answer probabilities suddenly collapse
- The Broader Spectrum of In-Context Learning — Argues in-context learning covers every behavioral adaptation a model makes that is not a weight change
- Forking Fast (recent paper) — Uses a statistical model of answer distributions to estimate uncertainty more cheaply than full resampling
- Stolen chain of thought paper — Shows a model's reasoning using ordinary words like musicals to stand in for unrelated, less transparent content
- Othello-GPT / Othello-Mamba research — Found similar interpretable internal representations across two different model architectures trained on the same game
Mentioned
Eric Bigelow · Goodfire · Bronson Shane · Apollo Research · Neel Nanda · Andrew Lambdan · Kimi K3 · DeepSeek R1 · Janus · Loom · Silico · Qwen












