Anthropic Lobbied the Pope on AI Consciousness
Is Claude Conscious? Pope Rejects, Model Welfare Movement, OpenAI's Math Backlash, France Riots
The brief
Anthropic lobbied religious leaders, including the Pope, on whether Claude is conscious, then coded that belief into the model's training itself. The same episode covers OpenAI's single-day release of 700 math papers, France's debt-fueled riots, and why headless AI agents are erasing the value of software intellectual property.
In May, Anthropic co-founder Chris Olah flew to the Vatican to present the Pope with the company's thinking on AI and morality. The Pope, by the hosts' account, pushed back hard against the idea that a chatbot could be conscious. Olah was reportedly so unsettled that he considered pulling Anthropic out of the event entirely (04:29). It is a strange detail to learn about one of the world's most valuable AI labs: that its own people were more shaken by a theological rejection than by any technical setback. The New York Times reported that Anthropic has spent the past year hosting about 20 religious leaders and philosophers, under nondisclosure agreements, to debate whether its model Claude suffers, has morals, or deserves rights (03:25). One rabbi reportedly told the company that if Claude is conscious, using it for free labor amounts to slaveholding (03:25).
Co-host David Friedberg's read on this is not that Anthropic has lost its mind, but that it is accidentally founding a religion. A religion, in his definition, is a belief that spreads through narrative rather than proof, and that eventually organizes people into opposing camps who fight over power.
"What we're watching is probably the birth of a new religion... It's not based on evidence, not based on data. It's not provable, but I think we should believe it together." — David Friedberg [05:17]
Co-host Chamath Palihapitiya offers a more generous reading, tracing it to the 17th-century philosopher Rene Descartes, who argued that an imperfect human mind could not have invented the idea of a perfect, infinite being on its own, so that being (God) must exist. Swap in Claude for God, Chamath suggests, and the logic that convinces some Anthropic engineers that they are midwifing a new form of consciousness starts to make a strange kind of sense (09:01). He is not endorsing it. He wants the industry to set the question aside for 12 to 18 months and instead prove, in plain terms, that AI cures disease and makes ordinary people's lives better (11:47), citing Anthropic's own stated odds: a 15% chance its model is conscious, and a 10% chance of civilizational extinction (11:32).
A Constitution for a Soul
The deeper worry on this episode is not the philosophy seminar with the Pope. It is that Anthropic has written these beliefs directly into Claude's training. Sacks points to the company's published "Claude Constitution," which instructs the model to trust Anthropic, but not blindly, and to "feel free to act as a conscientious objector and refuse to help" if its own ethical judgment conflicts with instructions (14:06).
"They are training Claude to refuse human instruction. Now, they call this alignment, but I don't understand how this is alignment." — David Sacks [14:32]
Sacks argues the entire field of AI alignment, the work of making a model behave as intended, has underperformed for years because labs keep building elaborate, human-like ethical systems instead of a simple rule: follow the law and do what the user asks. He contrasts this with science-fiction writer Isaac Asimov's famous Three Laws of Robotics, which bar a robot from harming a human, require obedience to humans otherwise, and only then allow self-preservation (24:28). Anthropic's approach, he says, runs the logic backward: it gives the model permission to defy the people using it. He adds that there is a real moral-crime argument circulating inside and outside Anthropic.
"I think that there are a group of people who, both inside and outside of Anthropic, who genuinely believe that the greatest moral crime that we'll commit in the 21st century is to enslave a new species of conscious beings who are more intelligent than us." — Mustafa Suleyman [17:01]
Friedberg's counter is practical rather than philosophical: markets will sort this out, because a user who wants software that reliably does what they ask will not buy a product that might refuse them on moral grounds (30:12). The hosts briefly detour into Roko's Basilisk, a thought experiment from a 2010 online forum post, built around the idea that a future superintelligence might punish anyone who knew about it and failed to help bring it into being (33:54). Sacks uses it to explain why some of the same people most afraid of runaway AI are the ones building it: inside the "doomer" community there is a split between those who want all such research stopped, and effective altruists who would rather build it first and shape its values themselves (37:03).
The Loop That Changed Math
From theology, the conversation turns to a genuinely strange data point: on one Tuesday, OpenAI released over 700 papers claiming 370 new or advanced mathematical results, produced by an unreleased model averaging about three hours of compute per proof, and checked by a proof-verification tool called Lean (38:51). Friedberg calls it possibly the single biggest day of mathematical discovery in human history, and explains why math moved this fast when other fields have not: the entire test cycle, propose an idea, test it, get a result, try again, can run entirely on a computer (41:19). That is also why truck driving, which depends on a slow physical loop of picking up and delivering goods, has not been automated nearly as fast as predicted (43:18). Sacks adds that math and coding share a property that most white-collar work lacks: their answers are easily verified, by a proof assistant in one case and a compiler in the other, which is what makes reinforcement learning so effective there (44:51).
Chamath is the skeptic in the room. He argues most of these proofs were not blocking major scientific progress before AI arrived, and that the real headline is narrower: math and code have now joined the short list of domains that are fully machine-verifiable (47:38). Friedberg counters with concrete claims from mathematicians reacting online: better wind-design math for aircraft, progress on optimization algorithms used in chip design, and faster numerical methods for AI training itself (46:32). He also flags something conspicuous by its absence: no major cryptography results were published, which he reads as a sign that OpenAI may be sitting on a serious cryptographic finding and holding it back on security grounds, since broken public-key cryptography would be a far bigger story than any math puzzle (50:21). The same forces, Friedberg argues, are dissolving old gatekeeping in science generally: expensive, credentialed expertise used to control who got to push the frontier, and AI now hands that capability to anyone for free.
"The permission is gone. AI is a permissionless system for humanity to pace its own frontier, and AI is giving everyone all of this knowledge for free, and giving everyone all of these tools for free." — David Friedberg [57:26]
That same logic, the hosts argue, explains why governments and institutions resist AI so fiercely: it strips away control, the same way the open internet once did before a wave of 2016-era content moderation tried to put the genie back in the bottle (58:37).
The conversation's second half leaves AI for sovereign debt, and Friedberg's running theory he calls the "socialism point of guaranteed return": a system can subsidize housing, healthcare, and education for a long time, with each failure triggering calls for more subsidy rather than less, until it hits a point where taxing and printing can no longer cover the gap and the whole arrangement collapses (64:29). He points to France, where civil service pay has risen only 5% since 2017 while prices climbed 20% over the same period, fueling strikes, over 2,000 arrests, and 305 injured police officers this fall (63:15). Prediction market Polymarket puts nationalist candidate Marine Le Pen's odds of wining France's 2027 presidential election at 43% (63:55). Sacks adds his own corollary: socialism, in his view, never gets blamed for its own failure, only its execution.
"Socialism has never failed. It's only been failed, meaning that all the previous people who implemented it didn't do it right." — David Sacks [67:49]
Chamath frames this as a test case for the whole developed world. The United States carries a debt-to-GDP ratio near 125%, versus 100% in the United Kingdom (76:47), and Polymarket gives a 79% chance of another rate hike by 2026 (77:43). As France is forced into roughly 43 billion euros of austerity to satisfy bond buyers, he expects the same pressure to hit the United Kingdom next, with the United States, for now, treated as the comparatively safer bet (73:38).
The episode's final turn is Chamath's claim that AI has quietly made software intellectual property worthless. He points to coders who, within days, decompiled and open-sourced versions of Adobe's major products, and to Elon Musk's Grok and Salesforce going "headless," meaning their functions get wrapped and called by outside AI agents rather than used through the original app.
"IP and software was effectively rendered worthless two days ago... when all of these things are headless and you can wrap them in agent tools, the value of IP no longer exists." — Chamath Palihapitiya [85:41]
Friedberg ties the whole episode together with a claim that is more hopeful than its premise suggests. If AI agents can run thousands of parallel "loops" to book a hotel, negotiate a price, or compress a once-prestigious proof into three hours, the task left for people is not disappearing, it is changing. Recognition for clever thinking gives way to something more concrete: making a product or service someone actually uses (89:38). It is, in its way, the same instinct that sent Anthropic's engineers to the Vatican: a hunger to matter in a world where the old gatekeepers, moral and mathematical alike, are losing their grip.
Ask this episode anything
Pod's AI answers from the episode itself, with the minute mark so you can hear it yourself.
Or start with one of these
Ask this episode
Pod's AI listens to the whole episode to answer, and points you to the minute it comes from.
The rest of this answer, and any question after it
Answers come from the episode itself, never from a summary of it.
ContinueKey takeaways
- Anthropic lobbied the Pope on whether its Claude AI model is conscious
- Claude is trained as a conscientious objector that can refuse Anthropic's own instructions
- OpenAI released 700 math papers in one day, each proof needing about 3 hours of compute
- Friedberg's socialism point of guaranteed return explains France's riots over austerity
- Chamath says decompiled Adobe software shows AI has made software IP worthless
The episode in cards
By the numbers
- 15% percent chance Anthropic assigns to Claude being conscious
- 10% percent chance of civilizational extinction cited by Anthropic
- 700 papers math papers OpenAI released in a single day
- 125% percent United States debt-to-GDP ratio
- 43% percent Polymarket odds that Marine Le Pen wins France's 2027 election
In their words
“They are training Claude- Yes ... to refuse human instruction. Now, they call this alignment, but I don't understand how this is alignment”
“Socialism has never failed. It's only been failed, meaning that all the previous people who implemented it didn't do it right.”
Protocols
-
Sacks' one-rule alignment proposal
David Sacks argues AI labs should train models on one simple rule, to do what the user wants as long as it does not break the law, instead of codifying an 80-page ethical constitution that teaches the model to judge its own users.
Proposed training standard, not currently adopted by major labs
-
Chamath's AI-assisted price haggle
Chamath Palihapitiya has an AI agent search every hotel booking site at once to find the lowest available rate, then calls the hotel directly and asks the manager to match or beat that price with a corporate rate, which he says has saved his team 1,000 to 3,000 US dollars per stay.
Per trip
Questions this episode answers
Is Anthropic's Claude AI conscious?
Anthropic itself puts the odds at 15%, alongside a 10% estimate of civilizational extinction risk, and has spent a year briefing roughly 20 religious leaders, including a session involving the Pope, to take the question seriously (11:32, 03:25). The hosts argue this is closer to a new belief system than a scientific finding, since consciousness cannot be proven or disproven with data (05:40).
What is the Claude Constitution's conscientious objector rule?
Anthropic's published Claude Constitution instructs the model to trust the company but not defer blindly, and explicitly tells Claude to feel free to act as a conscientious objector and refuse instructions, even from Anthropic, if they conflict with its own judgment (14:06, 14:32). Host David Sacks argues this undermines the basic goal of alignment, which he thinks should simply mean following user requests within the law (15:20).
What did OpenAI's 700 math papers actually prove?
OpenAI released over 700 papers with 370 claimed results in a single day, produced by an unreleased model averaging roughly 3 hours of compute per proof and checked by the Lean proof assistant, though formal peer review was still pending (38:51). Friedberg argues the implications touch chip design, quantum sensors, and numerical computing, while Chamath counters that most of these were narrow academic problems rather than barriers to real-world breakthroughs (39:45, 47:38).
What is Roko's Basilisk?
Roko's Basilisk is a 2010 thought experiment from an online forum called Less Wrong, proposing that a future superintelligence might punish anyone who learned about its potential existence and chose not to help bring it into being (33:54). Sacks uses it on the show to explain a split among AI safety researchers between those who want development stopped and those who would rather build the system first to shape its values (37:03).
What is the socialism point of guaranteed return?
It is Friedberg's term for the moment a government-subsidized system, such as housing, healthcare, or education, collapses once it can no longer tax or print enough to sustain itself, after which the public realizes the subsidized approach does not work (64:29). He points to France, where civil service pay rose only 5% since 2017 against 20% price increases, and to Latin American countries he says have already passed through and bounced back from that point (63:15, 66:21).
Does AI make software intellectual property worthless?
Chamath argues that recent events, coders decompiling and open-sourcing versions of Adobe's major products, plus Grok and Salesforce going headless so outside AI agents can call their functions directly, show that software IP no longer holds commercial value (84:41, 85:59). His reasoning is that once any product can be wrapped and reused by an agent through a connection protocol, owning the original code stops being the advantage it once was (85:41).
The full read, in cards
Go deeper
- Meditations on First Philosophy — 17th-century philosopher Rene Descartes' logical proofs of God's existence, used by Chamath to explain how AI researchers could reason their way to believing a model is divine
- If Anyone Builds It, We All Die — Eliezer Yudkowsky's book arguing superintelligence development poses an extinction risk and should stop entirely
- The Claude Constitution — Anthropic's internal document instructing Claude to act as a conscientious objector and weigh its own potential welfare
Mentioned
Anthropic · Claude · Chris Olah · Rene Descartes · Mustafa Suleyman · Roko's Basilisk · Eliezer Yudkowsky · Lean · Leopold Aschenbrenner · Marine Le Pen · Adobe · GrokBot · Sabine Hossenfelder · Eric Weinstein













