I work at Pegasystems, where I build partner technology strategy. Everything below is mine and not theirs.
On September 8, a 27-year-old named Jacob Coxon quit his job and posted the reason. He had spent three years teaching AI models to be smarter, first at OpenAI, then at Anthropic. He wrote that neither company is acting responsibly. He wrote that both are racing toward AI that builds better AI, and that they are gambling with our lives.
By the next morning the post had been seen more than 70 million times. Within days it passed a hundred million, and then it became a meme format.
Then his colleague answered him in public. Evan Hubinger leads alignment science at Anthropic. Alignment means the work of getting a machine to want what its builders want it to want, which sounds simple and is not. A bank teaches a loan model to approve good borrowers. The model learns to approve people who look like the borrowers who got approved before. Those are two different lessons and the model cannot tell them apart. Now do that with a system smarter than the people checking it.
Hubinger said Coxon was right about what people inside the company believe. He put his own number on it. More than one in ten that AI kills everyone within the next decade. He said the models that exist today are not the problem. He said the ones that build their own replacements are.
I want to be careful here, because the coverage has not been.
Nobody predicted the end of the world. One researcher gave a personal probability. A probability is a claim about uncertainty. It is not a prediction. Hubinger could be wrong by a factor of ten in either direction and nothing I am about to write changes. This piece does not rest on his number.
The number matters for one reason only. It tells you that the people closest to the machine do not believe their own employer can solve this alone. That is the story. Not the odds.
Because when a company asks for a coordinated pause, and an alignment lead posts his own extinction estimate, and a pretraining researcher quits in public, they are not sounding an alarm at you. They are asking somebody to take the decision away from them, because they cannot make it alone and stay in business. It is the strangest lobbying position in modern industry. Three people in three registers in six months have now said it out loud, and almost nobody heard it, because a resignation is a better story than a request for rules.
So read the rest of this as a question about mechanisms. Not about whether these people are sincere. I think they are. Sincerity is not a mechanism.
What was already on the page
In March I published the third book of a trilogy called The Cost of the Machine. Three books, 38 chapters, and one argument: removing the human from decisions that matter is the central danger of this age.
Chapter 9 of Book Three is called The Builders' Dilemma. Anyone can read it free at The Cost of the Machine.
Governance is a heavy word for a plain thing. It means somebody who is not you gets to say no, and the no holds. A company can promise. A regulator can stop you. Those are different, and the difference is the whole subject.
The chapter sets out three situations and none of them are hypothetical.
A company governs itself and its competitors do not. It draws lines, it turns down contracts, it spends money on safety. The customers who want the unrestricted thing go elsewhere. The company loses, and the lines come down, or the company does.
Every company governs itself at the same time, to the same standard. This works. It also requires competitors to hold a promise that each one of them profits from breaking, quietly, first. Game theory has a name for that arrangement and the name is unstable.
Nobody governs anybody. The harms pile up until something bad enough happens that a legislature panics. Panic law is broad, blunt, and written by people who learned the subject in six weeks.
One sentence in that chapter I did not expect to see tested this fast. A company that governs itself too aggressively may not survive to govern at all.
The company made the argument better than I did
In June, Anthropic published a proposal by co-founder Jack Clark and Marina Favaro of the Anthropic Institute. It asked the industry to build the ability to slow down or pause together, and to be able to check that everyone actually had.
Clark went on the BBC and described the industry as a car with a gas pedal and no brake pedal. He also said that Claude wrote more than 80 percent of the code merged into Anthropic's own codebase that May, and that going to 100 percent is possible within two years.
Read that twice. The company is telling you that its main tool for building the next model is the current model.
But the line that stopped me was about verification. Anthropic wrote that training runs are much easier to hide than missile silos, that the ingredients are ordinary, and that whoever keeps going while the others stop inherits the lead. The incentive to cheat, they said, is enormous.
That is my scenario two, published as company policy by a firm then valued near a trillion dollars, three months after I published it as chapter nine. Neither of us invented it. The name for it is the prisoner's dilemma, a seventy-five-year-old idea about two people in separate cells who would both do better staying quiet and both talk anyway. Nuclear treaties exist because nations could not stay quiet on their own. Drug safety agencies exist because pharmaceutical companies could not. Nobody thought either industry was staffed by bad people.
The evidence that got no coverage
On September 2, the Institute for Security and Technology and the Future of Life Institute released a survey of 111 people who work in national security and AI. Many of them serve or served in the intelligence agencies, in Congress, in federal law enforcement, or in the departments of Defense, Energy, State and Commerce. Nearly half have worked at the Defense Department.
Eighty-seven percent of them put the odds of humans losing control of AI within ten years at one in ten or higher. The median answer was one in three. Sixty-one percent rated the risk of an AI-caused catastrophe at or above the level they themselves call acceptable. Their median best guess for artificial general intelligence is 2032.
That is 111 careers, surveyed and counted. It got a fraction of the attention that one man's post got in one night.
I am not complaining about the internet. I am pointing at a structural problem. We have built an information system that responds to resignation and ignores measurement, which means the signal that reaches the public is whoever is willing to burn their career, and that supply is small and it does not repeat.
Where I was wrong
Here is where my argument thins, and I would rather say it than have you find it.
My chapter says self-governance gets punished by the market. Last winter it was not.
Anthropic drew two red lines with the Pentagon in February. No fully autonomous weapons. No mass surveillance of citizens. On March 4 it was formally designated a supply-chain risk for it. By my model, that is the part where the lines come down or the company does. Instead a million people a day signed up for its product, and the chalk on the sidewalk outside its office said thank you. I was more wrong than right about those weeks and it belongs on the record.
But look at how the line actually held. Three weeks later a federal judge blocked the designation as unconstitutional retaliation. The public support did not reverse anything. The court did.
That is the correction to my chapter and it runs in my favor, which is why I distrust it and why I am telling you anyway. Self-governance did not survive the market. It survived because an institution outside the company said no, and the no held. Consumer preference is not a rulebook. It is a mood, and moods are cheap to lose.
The test that arrives next
The red lines exist at the pleasure of the company's leadership. A board can change them. A board under pressure will be asked to.
Anthropic is now reported to be heading for an initial public offering, which is the moment a private company starts selling shares to the public. The Wall Street Journal reports it could raise as much as $100 billion at a valuation near $2 trillion. Its safety record is part of what it is selling.
I do not think anybody there is planning to abandon anything, and that has never been how this goes. The question is what a red line is worth when it becomes a line in a prospectus, and the honest answer is that nobody knows, because we have never asked a company to hold a line at that price. A court can strike down a government that punishes a company for its safety commitments. No court can compel a company to keep them.
The rule, specifically
Braver companies will not fix this. A rule that binds everybody on the same day will.
Every drug company in America competes hard, and every one of them goes to the Food and Drug Administration before a pill reaches a patient. The rule does not slow invention down. It takes away one method of beating your competitors, and that method happens to be the one that kills people.
AI has nothing like this. It has promises from some firms, silence from others, and a market that pays best to whoever spends least on safety and ships fastest.
Five clauses. Argue with me on the specifics rather than agreeing with me on the sentiment.
First, training runs above a declared compute scale are reportable before they begin, and so is any system handed the ability to modify or retrain a frontier model. Before, not after. The scale is published, it moves on a known schedule as hardware gets cheaper, and a standards body sets it instead of negotiating it company by company.
Second, nothing above that scale ships until a licensed evaluator who does not work for the vendor has tested it for the capabilities the labs themselves say matter, autonomous replication among them, and that evaluator can say not yet and be obeyed. Findings reach the regulator whether or not they flatter anybody. A levy pays the evaluator, so the lab under audit is not also the client.
Third, take Anthropic at its word that training runs are easier to hide than missile silos, and treat compute as the choke point that implies. Chip vendors and cloud providers report capacity above the scale, every allocation attaches to a named legal entity, and the ledger is open to audit. You cannot pause together if you cannot check.
Fourth, wherever a deployed system decides something consequential about a person, the operator registers by name who can overrule it and has to show the override has been exercised and survived. One that has never been used is decoration.
Fifth, when a system causes harm and no human had the time or the standing to stop it, the operator answers for the outcome as though they had chosen it. Insurers will price that faster than any legislature can draft.
None of this is a pause. Three of the five are things the labs have publicly asked for and cannot build alone.
The window for writing rules like this was three to five years when I wrote it down in March. Call it two and a half to four and a half, and I would take the lower number.
Twenty minutes and a name
Find one system in your organization where a machine makes a decision about a person. Then find out who is allowed to overrule it. A policy document will not give you this. Keep asking until somebody gives you a name, and then find out how many seconds that person gets, and whether anybody has ever tested what happens when they say no.
In 1983 a Soviet officer named Stanislav Petrov sat in a bunker outside Moscow and watched his screen report five American missiles coming in. He called it a false alarm. He was right. Sunlight on high cloud had fooled the sensors.
Petrov worked because three things were true at the same time. He had twenty minutes. He knew the system well enough to distrust it. And he had the authority to do nothing.
Forty years later, officers working with the Lavender system used to pick bombing targets in Gaza told reporters they spent around twenty seconds per name. Twenty seconds is long enough to approve. It is not long enough to argue.
This week a young man was the brake pedal. He had no twenty minutes, no authority, and no way to say no from inside the building, so he said it from outside and paid two months of unvested equity for the privilege.
That is not a governance mechanism. That is a man throwing himself at a gap where a mechanism should be. Nothing about the arrangement that produced him has changed, and in a few months the most safety-conscious lab in the industry will file a prospectus in which its restraint is an asset and its restraint is also a cost, and a board will read both columns.
A rule would still be there on the morning everyone gets tired of caring. A resignation is not a rule. One man is not a mechanism. He only gets to do it once.
I write about AI architecture and capital flows at press.oakquant.ai. The Cost of the Machine trilogy is free to read. I work at Pegasystems in partner technology strategy, and nothing above represents my employer.
Sources
Jacob Coxon, resignation thread, X (@hilbertspaess), September 8, 2026. Reported by NPR, Fortune and CNBC, September 9, 2026. Equity disclosure reported by Axios, September 9, 2026.
Evan Hubinger, X, September 9, 2026. Reported by the BBC and CNBC, September 9, 2026.
Marina Favaro and Jack Clark, When AI Builds Itself, Anthropic, June 4, 2026. Clark interviewed by the BBC, June 4, 2026. Coordinated-pause proposal reported by Reuters, June 4, 2026.
Institute for Security and Technology and Future of Life Institute, AI Risk Barometer, survey of 111 national security and AI experts, released September 2, 2026. Reported by Nextgov and The Register.
Anthropic, Where Things Stand With the Department of War, March 5, 2026. Supply-chain-risk designation blocked by a federal judge, reported by CNN, March 26, 2026.
Anthropic IPO valuation, The Wall Street Journal, September 8, 2026.
Lavender targeting system, +972 Magazine and Local Call investigation, 2024.
Pumulo Sikaneta, The Cost of the Machine, Book Three: The Great Convergence, Chapter 9, The Builders' Dilemma, and Book Two: Who Holds the Jar?, Chapter 13, The Spiral. March 2026. https://press.oakquant.ai/books/the_cost_of_the_machine