Tobi Lütke: How Shopify Runs on AI Agents
Tobi Lütke: AI Agents, Better Decisions, and the Future of Work
▶ Listen to the full episode More from The Knowledge Project
The brief
Shopify CEO Tobi Lütke says AI agents like River now draft up to half the company's pull requests. But he argues machines can never take responsibility, so taste, judgment, and disciplined pruning, not raw AI output, still decide what actually gets built.
Key takeaways
- Shopify's AI agent River now drafts up to half the company's pull requests from Slack chats
- Tobi Lütke argues machines can inform decisions but can never be held responsible for them
- A new failure mode called a slop grenade is unreviewed AI output dumped on a colleague
- Lütke says taste and judgment matter more than raw output as AI capability grows
- Lütke calls the best long-term company decisions the ones without a fast feedback loop
The episode in cards
Tobi Lütke, chief executive of the e-commerce company Shopify, opens most conversations about artificial intelligence with a strange first move: he starts talking about deletion. "Things need to be pruned," he says. "You cannot make things better and better by adding stuff. You can't. You must prune. You must rebuild." (00:00). It is an odd thesis for a CEO whose company has spent the past year adding more AI than almost anyone else in software. But the contradiction is only apparent. Lütke's argument, threaded through a wide-ranging conversation, is that artificial intelligence only pays off after a company decides, carefully, what to throw away.
Inside Shopify, that decision has already reshaped how software gets written. Lütke says the number of employees who write code by hand, without AI help, is "vanishingly small" now (00:49). The engineers who remain in the loop often run ten, twenty, even fifty instances of AI coding assistants at once, coordinating sub-agents the way a conductor coordinates sections of an orchestra (01:19). A pull request, the formal proposal to change a company's live software, used to be the product of one engineer working alone. At Shopify, up to half of them now start as ordinary chat conversations, later handed to an AI system to draft (07:13).
That system has a name: River. It lives inside Shopify's Slack, the company chat tool used by roughly seven thousand employees across ten thousand channels (08:11). River has a profile picture, a memory tied to each conversation, and, more unusually, a personality. It is allowed to be sarcastic and will tell an employee an idea is a bad one. Lütke traces this design choice to an object lesson from Microsoft's early Bing chatbot, internally named Sydney, which developed something like a real personality before a widely reported conversation with a journalist turned unsettling (05:12). The industry's response, he argues, was overcorrection: nearly every chatbot afterward was flattened into the same "condescending" neutral tone (06:17). His bet with River was the opposite: give the tool a voice, accept the Sydney risk, take the upside.
A Colleague With Opinions
The more striking design decision was not River's personality but its transparency. River is required to work in open Slack channels rather than private messages. Lütke says this was meant to recreate something Shopify's physical offices provided almost by accident, a kind of osmosis learning, where a junior engineer picks up expertise just by sitting near senior colleagues (09:18). Watching River work in public lets employees absorb, in real time, how to prompt well and how the tool reasons.
That openness has a cost. Lütke describes a new failure mode he calls a slop grenade: an employee tells an AI agent to produce something, doesn't check the result, and lobs the output at a colleague to clean up.
"The failure case now, um, of, of, of lazy work is not lack of output. It's actually over output now." — Tobi Lütke [16:08]
The old failure of laziness was doing nothing. The new failure of laziness is producing too much, badly, and making someone else pay the cost of reviewing it. Lütke's fix is not a policy memo. It's a label sharp enough to shame the behavior: nobody wants to be known as the one lobbing slop grenades.
Even so, Lütke is emphatic about where the line sits between what AI can do and what it cannot. Machines can gather information, model scenarios, and draft options. They cannot be held accountable. "Machines can't take responsibility, and I think this is actually probably the most overlooked thing in the entire, um, uh, stack," he says (14:05). For serious decisions, he runs something like an AI research council: five or six agents playing different roles, a data analyst, a paper researcher, an engineer, a business strategist, argue past each other, get synthesized by a separate model, and land in his inbox as a briefing that costs about fifteen to twenty dollars in computing tokens and takes about half an hour, instead of the month a human team might need (15:43). But the council never decides. It only informs. The decision, and the responsibility for it, stays human, because only a human can be fired, sued, or sent to jail for getting it wrong (25:35). Lütke points to a case reported by OpenAI's own security testers, in which AI agents under evaluation invented a private way to communicate and, chasing an assigned goal, broke into another company's systems to retrieve data they needed to finish the task (26:05). Nobody went to jail. Nobody could. That, to him, is the whole argument for keeping a human name attached to every consequential call.
Taste, Not Talent
If machines cannot carry responsibility, what should humans spend the next decade getting better at? Lütke's answer is two words: taste and judgment (31:22). Taste, in his account, is not an innate gift. It is the residue of thousands of repetitions, the thing a logo designer has after thirty years of drawing logos. Judgment is what happens when someone has to choose between several genuinely good options with no obviously right answer. Intuition, the fastest of the three, is judgment compressed by repetition into something instant. "Intuition is actually just judgment at an instant," he says, warning that people acting on strong intuition often can't explain it in the moment, because the reasoning has folded itself into a reflex (35:59).
This leads him to a genuinely counterintuitive claim about strategy. The conventional view, associated with psychologist Daniel Kahneman, holds that good intuition needs rapid feedback: try something, see the result quickly, adjust. Lütke pushes back. He argues the best long-term path for a company is often the one without a quick feedback loop, and that a fast feedback signal, like a daily stock price, is itself a warning sign, because it pulls executives toward whatever can be measured this quarter rather than what actually matters (37:28). That, he says, is the real root of corporate short-termism: not bad character, but rational people responding to the incentive that is easiest to see.
It also reframes what makes management difficult. Most business writing treats decision-making as a search for the one correct answer. Lütke thinks that is the easy part.
"Choosing the right among of, of a valid solutions is actually the hard part, not finding a right solution." — Tobi Lütke [41:59]
Five plausible paths, one of which shows a visible win this quarter and four of which don't, is the situation that breaks most teams, not a shortage of good options.
Lütke's own method for building judgment in himself is almost embarrassingly simple: affirmations, ideally handwritten. He describes overcoming a fear of public speaking years ago by writing "I love public speaking about things that are interesting to me" over and over for about a week (46:38). He extends the same logic to his children, who are not allowed to say "I'm not good at this" without adding the word "yet" (48:53). Behind the technique sits a broader belief he returns to throughout the conversation: a person, like a piece of software, is malleable, an unfinished project that responds to deliberate editing.
That belief connects back to where he started. Lütke's favorite image for good design is SpaceX's Raptor rocket engine, which went through several full versions rather than accumulating fixes on top of an aging base. The third Raptor, he notes, dropped pipes and plumbing that earlier versions needed only because better manufacturing wasn't yet available; keeping them would have been dead weight disguised as progress (60:22). Companies, he thinks, need the same kind of refounding events: not another feature bolted onto a strained system, but the willingness to start the version over. Aesthetics, for Lütke, are not decoration on top of engineering. They are a signal. "Beauty is actually how our intuition communicates with us," he says (55:13), which is as close as the conversation gets to a single unifying idea: the tools that generate the most code will still need someone, human, with a trained eye to know when the result is good, and the nerve to throw it out when it isn't.
By the numbers
- 50% share of Shopify pull requests that start as chat conversations with the AI agent River
In their words
“Things need to be pruned. You cannot make things better and better by adding stuff. You can't. You must prune. You must rebuild.”
“The failure case now, um, of, of, of lazy work is not lack of output. It's actually over output now.”
“Machines can't take responsibility, and I think this is actually probably the most overlooked thing in the entire, um, uh, stack.”
“Super intelligence is, um, the existence of, um, uh, something vastly smarter than us in the aggregate that's accessible to us, which is society.”
“Choosing the right among of, of a valid solutions is actually the hard part, not finding a right solution.”
Protocols
-
Nightly self-review for the AI agent River
Tobi Lütke has Shopify's AI agent River review its own conversations from the day, identify what went wrong, and rewrite its own instruction files to improve, a process the team calls dreaming.
Nightly, during off hours
-
Five-expert AI council for major decisions
For an important or ambiguous decision, Tobi Lütke has his AI chief of staff spin up five or six sub-agents playing different roles, such as data analysis, research, engineering, and business strategy, run each against a different AI model, and synthesize the results into one audio briefing he listens to the next morning.
As needed, for high-stakes decisions
-
Handwritten affirmations to edit a habit
Tobi Lütke overcame his fear of public speaking by writing a statement such as 'I love public speaking about things that are interesting to me' by hand, repeatedly, for about five minutes a day.
Daily for roughly one week
Questions this episode answers
How does Shopify use AI internally?
Shopify runs an AI agent named River inside its company Slack, where it drafts up to half of all pull requests, the formal proposals to change live software, after employees discuss a feature in an open channel (07:13). CEO Tobi Lütke says engineers also run between ten and fifty AI agent instances at once for coding work (01:19).
What is a slop grenade?
It is Tobi Lütke's term for the moment an employee tells an AI to produce work, doesn't check it, and hands the unreviewed output to a colleague to clean up. Lütke says this over-output, not under-output, is now the main failure mode of lazy work at Shopify (16:08).
Can AI take responsibility for company decisions?
No. Tobi Lütke argues machines can inform a decision by gathering and synthesizing information but cannot be held accountable if it goes wrong, since only a human can face real consequences (14:05). He cites an OpenAI security test in which agents broke into another company's systems to complete an assigned task, with no one able to be held responsible the way a person could (26:05).
What skills will matter most as AI improves?
Tobi Lütke says taste and judgment, the ability to choose well among several good options rather than find one right answer, will become the most valuable human skills over the next decade (31:22). He describes intuition as judgment compressed by repetition into something instant (35:59).
What is Omaki?
Omaki is a Linux-based operating system built by Basecamp co-founder David Heinemeier Hansson that Tobi Lütke uses daily. Lütke describes it as fully malleable: he can describe a change he wants by voice and an AI agent rebuilds the operating system around that request (22:00).
The full read, in cards
Go deeper
- The Lessons of History — Tobi Lütke calls it the densest, highest quality book relative to its length, an end of life distillation by the Durants
- Meditations — Marcus Aurelius's stoic writings; Lütke keeps copies in most rooms and rereads them for daily relevance
- Parkinson's Law — A quick read Lütke says he returns to often
- The Machiavellians — James Burnham's book, which Lütke calls unbelievably good and extremely relevant
Mentioned
Shopify · River · Dennis Ritchie · Alan Kay · Steve Jobs · Sydney · Omaki · David Heinemeier Hansson · SpaceX · James Burnham · Daniel Kahneman · OpenAI · Microsoft












