The last temptation is the greatest treason: to do the right deed for the wrong reason.
On June 4, 2026, Anthropic — one of the handful of companies building the most capable artificial intelligence systems in the world — published an essay titled When AI builds itself. Its argument was, in essence, that the world should slow down. The authors, Marina Favaro and Jack Clark, wrote that it would be good for the world “to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology.”1
The essay was not vague. It disclosed internal data: as of May 2026, more than 80% of the code Anthropic merges into its own codebase is written by its model, Claude — up from low single digits before early 2025.2 It described the approach of recursive self-improvement, the point at which an AI system could autonomously design and develop its own successor. It was careful to say this has not happened and is not inevitable, “but could come sooner than most institutions are prepared for.”
This is a serious document, written by serious people, making a serious claim. The correct first response is to take it seriously.
And yet Eliot’s line will not leave me alone.
Eliot put those words into the mouth of Thomas Becket, Archbishop of Canterbury, awaiting murder in his own cathedral. Becket has refused every temptation — power, safety, compromise — and faces the last and subtlest one: to embrace martyrdom because it is glorious, to do the right thing for the reward of having done it. The deed is correct. The reason corrupts it.
The unsettling relevance is this. “Slow down” is not a neutral sentence. It means one thing from a regulator, another from a new entrant, and something else again from the acknowledged front-runner in the field. Anthropic is the front-runner: a recent funding round valued it at roughly $965 billion, and it has filed to go public.3 A market leader calling for an industry-wide pause may be expressing genuine alarm. It may also, by the same words, be proposing to freeze the race at the moment it is winning. The act is identical in both cases. Only the reason differs - and the reason changes everything.
There is a sharper version of this problem, and it is not hypothetical either. Four months before asking the world to slow down, the same company softened the central pledge of its own Responsible Scaling Policy — turning a categorical commitment to halt development when safety measures fell behind into a conditional one that applies only when Anthropic judges itself not to be losing its lead; in other words, a commitment to pause only while it is ahead.4 The company framed this as realism rather than retreat: a way to set safety targets it could actually meet. Perhaps. But it means the firm now urging the world to slow down had recently loosened the one rule that would have obliged it to slow down itself. I do not raise this to settle the question of motive. I raise it because it is exactly the kind of fact that the question of motive exists to notice.
This is not a hypothetical suspicion. The critique already exists, voiced from inside the industry: that Anthropic’s safety advocacy functions as “regulatory capture”, raising barriers that burden competitors and open-source developers more than Anthropic itself.5 You do not have to believe that critique to see the problem with it — the same problem that applies to its opposite. From the outside, we cannot confirm it, and we cannot dismiss it. We have the company’s words and its position in the market, and no access to whatever sits behind them.
Here is where the essay’s own logic deserves credit, because it does not ask us to. Favaro and Clark are explicit that a unilateral pause “accomplishes much less” than a coordinated one — that it “would change who the front-runner is, but it would not create the wider deliberative process that is currently missing.” They argue that any meaningful slowdown would require multiple frontier labs, in multiple countries, stopping under the same conditions, with each able to verify that the others have actually stopped. They concede that AI training runs are far easier to conceal than missile silos, which makes the verification problem harder than the Cold War arms-control regimes they invoke as precedent — including the Intermediate-Range Nuclear Forces Treaty.6
In other words, the essay’s strongest passages are the ones that move the question away from anyone’s motives and toward mechanism: not “trust us,” but “build the system that makes trust unnecessary.” That is the right instinct. When you cannot see inside a decision, what you can ask for is a way to check it from the outside.
So the temptation I want to name is not Anthropic’s. It is ours — the readers’.
There are two ways to fail here, and they are opposite.
The first is deference: to accept the right deed and never ask the reason, because the deed is reassuring and the people seem sincere. This is the failure of an audience that wants to be told the adults are handling it.
The second is the failure I find more tempting, because it flatters the person who falls for it: reflexive suspicion. To decide in advance that every call for restraint is a maneuver, every safety pledge a marketing document, every stated fear a moat under construction. It costs nothing and it explains everything, which is exactly why it should be distrusted. A suspicion that no evidence could ever overturn has stopped being an argument. And it carries a real cost: if no good deed can ever be credited, no one has any reason to attempt one.
Becket did not answer his last temptation by refusing the deed. He went through with it and let the question of his own motives stand unanswered. The deed was done; the reason stayed open. That is the harder discipline, and it is the one this publication intends to practice.
Not deference. Not suspicion. Scrutiny — which is neither, and harder than both. Scrutiny asks what mechanisms exist, what would prove the claim wrong, who checks the people doing the checking, what the incentives reward no matter what anyone intends. It can credit a good argument even when the person making it stands to gain. And it keeps asking the uncomfortable question without mistaking the question for an answer.
The AI industry is full of right deeds at the moment: safety research, alignment work, calls for coordination, essays urging the world to slow down. Some are exactly what they appear to be. Some are not. Most, if human beings are involved, are tangled — sincere and strategic at once, in proportions no one outside can cleanly measure, and perhaps no one inside can either.
We will not pretend to untangle them with certainty. We will refuse to stop asking. Because the last temptation — to take the right deed and never inquire after the reason — is one that belongs to the audience as much as the actor. And for a publication that exists to watch this, accepting it would be its own small treason.
Polanyi is an independent index of what technology does to us. This is the first letter.
1 Marina Favaro and Jack Clark, "When AI builds itself," The Anthropic Institute, published June 4, 2026. https://www.anthropic.com/institute/recursive-self-improvement. All quotations in this essay are taken directly from that piece.
2 Ibid. The essay specifies that the ">80%" figure measures the share of lines merged to production attributable to Claude, and notes this is more conservative than the "90% or more" some Anthropic leaders have estimated publicly, since the latter includes scripts and experimental code.
3 Anthropic confidentially filed a draft S-1 registration statement with the U.S. Securities and Exchange Commission on June 1, 2026. The filing followed a $65 billion Series H round in late May 2026 that set its valuation at roughly $965 billion — up from about $380 billion in February 2026 — and analysts widely expect a public debut in the trillion-dollar range. See CNBC ("Anthropic confidentially files IPO prospectus with SEC," June 1, 2026) and TechCrunch ("Anthropic files to go public," June 1, 2026).
4 Anthropic published Version 3.0 of its Responsible Scaling Policy on February 24, 2026. The 2023 version contained a categorical commitment not to train or deploy a system unless the company could guarantee adequate safety measures; the revision replaced this with a conditional commitment to "delay" training only under narrower circumstances — roughly, when the company judges both that catastrophic risk is significant and that it is not losing its competitive lead. Anthropic board member Holden Karnofsky defended the change as enabling more realistic safety targets, and METR's Chris Painter described it as a shift into "triage mode" as risk-assessment methods struggle to keep pace with capabilities. Critics read it as the removal of the policy's binding force. See TIME ("Exclusive: Anthropic Drops Flagship Safety Pledge," February 24, 2026) and Bloomberg (February 25, 2026).
5 The "regulatory capture" charge has been made publicly by David Sacks — venture capitalist and the Trump administration's AI and crypto policy lead — who has accused Anthropic of "a sophisticated regulatory capture strategy based on fear-mongering," arguing its policy posture would burden cheaper open-source models more than Anthropic itself. Anthropic rejects the characterization, and CEO Dario Amodei has publicly disputed it. See Semafor (October 17, 2025) and Fortune (June 5, 2026). The point here is not that the charge is correct, but that it cannot be settled from outside the company.
6 The Intermediate-Range Nuclear Forces (INF) Treaty was signed by the United States and the Soviet Union in 1987, banning ground-launched ballistic and cruise missiles with ranges of 500–5,500 km. The United States announced withdrawal in 2019 and the treaty lapsed the same year — a precedent worth holding in mind when invoking arms control as a model for durable restraint.
