On January 9 the Secretary of Defense signed a memorandum ordering the American military to become an “AI-first” fighting force. Six pages, addressed to senior Pentagon leadership, combatant commanders and defense agency directors, with a classified annex delivered by separate cover.
Buried on page five, under a heading about system architecture, sits this:
“In the AI arms race, system architectures must enable component replacement at commercial velocity to maintain overmatch.”
Note what that sentence assumes. The memo does not defend the arms race as a proposition; it treats the race as the weather, the settled condition inside which procurement decisions get made, and it says so in a directive rather than an op-ed or tweet. Pete Hegseth returns to the word race throughout. The transformation is a race fueled by commercial innovation from the American private sector. Military AI will be a race for the foreseeable future. The competition is dynamic and unpredictable.
The conclusions follow the way they always follow from that particular premise. Under the heading “Speed Wins,” the memo instructs the department to weaponize learning speed and to treat cycle time as a decisive variable, then states the tradeoff without ornament:
“We must accept that the risks of not moving fast enough outweigh the risks of imperfect alignment.”
There is a specified instrument for acting on that sentence. Under “Wartime Approach to Blockers,” the memo directs the department to eliminate obstacles across data sharing, authorizations to operate, testing, certification, contracting and hiring, and tells officials to handle risk tradeoffs and “equities” as questions to be settled as though under fire. The scare quotes around equities are the memo’s own, and they carry the whole attitude, since an equity is what a lawyer or an ethics officer or a privacy review would call a reason to slow down. To make the disposal routine, the memo establishes a monthly “Barrier Removal Board” with authority to waive non-statutory requirements and escalate the rest. Constraint is no longer something the department weighs case by case. It is a queue, worked down on a schedule, by a body whose reason to exist is clearing it.
Then, on page five, a heading that deserves to be read twice:
“Clarifying ‘Responsible AI’ at the DoW - Out with Utopian Idealism, In with Hard-Nosed Realism.”
The paragraph beneath it welds two instructions together. Diversity, equity and inclusion have no place in the department, it says, so the department must not employ models carrying ideological “tuning” that interferes with objectively truthful responses to user prompts. And in the same breath the department must use models free from usage policy constraints that may limit lawful military applications. Two directives follow. CDAO will establish benchmarks for model objectivity as a primary procurement criterion within ninety days. The Under Secretary for Acquisition and Sustainment will insert standard “any lawful use” language into every departmental contract procuring AI services within a hundred and eighty days.
Stripping safety constraints out of military AI contracts was not issued as a freestanding operational judgment about speed. It arrived as a single clause in a paragraph abolishing responsible AI as utopian idealism, bracketed with a culture-war instruction about model tuning, accompanied by an order to build a political-neutrality benchmark and buy against it. The constraints at issue are not decorative. They govern refusal behavior, the engineered unwillingness of a system to perform a task, and most of that behavior lives not in the weights but in the removable scaffolding around them, the system prompt, the wrapper, the classifier that reads a request before the model answers. That is what makes the demand so cheap to meet. A vendor told to deliver a model “free from usage policy constraints” is mostly being asked to remove a layer that was bolted on rather than trained in. Even the refusals trained deeper into the weights give way to a documented procedure called abliteration, which finds the internal direction that carries a model’s capacity to say no and suppresses it. Abliteration takes skill, but it takes no retraining run and no supercomputer, and that is the point. The contract clause reaches the guardrail because the guardrail was always the shallowest part of the system.
♦
Anthropic, the first frontier AI company on US government classified networks, held two lines that predated the memo: no mass domestic surveillance, no fully autonomous weapons. In July 2025 the Chief Digital and Artificial Intelligence Office had awarded the company a two-year prototype agreement with a $200 million ceiling, and its model was deployed on classified networks through Palantir’s Maven Smart System.
On February 27, after the company declined to drop the restrictions, Hegseth directed that it be designated a supply chain risk to national security, an authority built to keep Huawei and Kaspersky out of federal systems. President Trump ordered every federal agency to cease using its technology, and agencies began shedding contracts. Hours after the designation was announced, OpenAI announced its own Pentagon agreement, negotiated over a few days, which Sam Altman subsequently called rushed and conceded had looked opportunistic and sloppy. By March 2 he was posting amendments to the terms on X, driven not by legal review but by public backlash, including an open letter signed by hundreds of OpenAI and Google employees.
Two facts complete the sequence. The Washington Post reported that the excluded company’s model, running through Palantir on classified networks, was generating proposed targets in the Iran campaign, and that commanders had grown dependent enough that if the vendor ordered a halt, the administration would invoke government powers to retain access until a replacement existed. A federal exclusion order was in force against the maker of a system still nominating targets, and the government’s position was that it would compel continued access if the maker objected. Meanwhile Altman told an all-hands that his company does not get to make operational decisions, and named the mechanism himself. There would be at least one other actor, he assumed xAI, effectively offering to do whatever the government wanted.
The General Services Administration has since proposed extending the any-lawful-purpose requirement to civilian procurement, with a draft clause barring AI systems from refusing outputs or analyses on the basis of a contractor’s discretionary policies. What started as a Pentagon rule for targeting tools is migrating toward every federal desk.
Set aside what you think of any of these companies. A vendor’s refusal to sell an unrestricted targeting tool got processed as a threat to the national supply chain, and the word that made the processing feel natural was race.
♦
That vocabulary has a lineage, and that lineage is where it fails. On September 3, Senator Bernie Sanders and Representative Greg Casar announced the Ban Artificial Superintelligence Act, which would permanently prohibit developing or deploying superintelligent AI, freeze advanced development until a federal regulator writes safety rules, and commit the United States to pursuing international agreements, allied coordination and export controls to prevent anyone anywhere from building such a system. Trade coverage describes a maximum penalty of twenty years, matching unlawful nuclear weapons development, and an international architecture modeled on the Nonproliferation Treaty. In May, pressing for talks with Beijing, Sanders had invoked Reagan and Gorbachev directly.
Hegseth and Sanders are drawing on the same reservoir. One reaches for the corner of the history where American scientific supremacy is a survival requirement and domestic constraint a peacetime luxury, the other for the corner where scientists warn, statesmen negotiate and a treaty holds the line. Both are borrowing from 1949, and both inherit the same defect.
The nuclear control regime worked because fissile material is rivalrous and conserved. A kilogram of highly enriched uranium in your possession is a kilogram not in mine. Safeguards inspectors do arithmetic. Diversion leaves a hole in the ledger, and that hole is the evidence. Meanwhile model weights consist of a computer file. Theft involves copying, so the victim keeps full use of what was taken and may never learn it was taken. There is no inventory to reconcile because no quantity is conserved. The word “nonproliferation” carries an enforcement apparatus that runs on a property this asset lacks, and it has crossed the American political spectrum with nobody catching the discrepancy because the analogy pays out for everyone who reaches for it.
Buried in the Sanders bill is the exception that proves it. Export controls survive translation, because compute is rivalrous, countable, geographically fixed and already governed by a functioning regime. The reason turns on a distinction the bill never spells out but depends on completely: the difference between building a frontier model and running one. Training is a construction project visible from orbit, tens of thousands of serialized chips drawing a substation’s worth of power in one place for months. Running the finished model is almost nothing, a file that cost a gigawatt to make answering questions on a rack of gaming cards in a basement and throwing off no more heat than a space heater. The factory cannot be hidden yet every copy that leaves it can, which is why the serialized chip is the last object in this economy a state can actually count. Catch the silicon at manufacture and there is a chokepoint, but once the model exists there is nothing left with edges to inspect. The sole enforceable provision in a nonproliferation-framed bill governs the substrate rather than the artifact, and it only works because it abandons the nuclear analogy where it fails.
One regime does govern a threat shaped like this one, and it cuts against the argument so far. The Biological Weapons Convention has held for fifty years over an asset that is non-rivalrous, uncountable and made of information, namely the knowledge of how to build a pathogen. It has almost no verification, and what enforcement it has comes from stigma rather than counting. That leaves the claim standing but narrowed. Nonproliferation language can lend AI the half that runs on norm and none of the half that runs on the ledger. The borrowing runs into one more complication, and the biology makes it vivid. Pathogen knowledge is the thing the Convention tries to keep scarce, and a model that folds proteins to order generates it on demand.
The one arms-control regime whose logic transfers to AI governs a subject matter the technology is already busy mass-producing.
♦
Everything downstream of that error rests on a second problem. Nobody outside the labs can measure what is being raced over.
A question that ought to be simple: how far behind the closed frontier are the best open-weight models? Epoch AI, applying its composite capability index, puts the lag at an average of four months since January 2026, roughly eight index points, slightly wider than the three months it measured from 2023 to late 2025. OpenRouter, surveying the same field in June, reported a consistent three-to-six-month gap sustained for eighteen months and explicitly not widening. Arena’s evaluation data this month says the reverse, a twenty-nine-point Elo gap described as the widest since tracking began. Epoch notes the answer moves depending on whether you use public benchmarks, where the gap runs four to six months, or private ones, where it runs eight to ten.
Four defensible answers, all measuring capability on tests the labs help design and optimize against. Thinking of this as a measurement error ignores the finding: that spread is what an unmeasurable contest looks like from outside.
In an opaque contest, threat estimates are endogenous to the interests of whoever produces them, and the primary sources of capability claims about the American labs, and about their Chinese rivals, are the American labs and the people who fund them. This is not a conspiracy theory, but a recognized failure in the literature. Roberta Wohlstetter’s study of Pearl Harbor showed that signal and noise are qualities the receiver assigns, sorted by what the receiver already expects and already wants; the warning that gets heard is the one that flatters a standing commitment. No competent intelligence service would accept the labs’ self-report in any other domain for the same reason no one lets the subject grade its own paper.
Hegseth’s own memo concedes the measurement problem from the other direction. Under “AI Model Parity” it directs CDAO to establish a vendor cadence enabling the latest models to be fielded within thirty days of public release, a primary procurement criterion. A department that binds itself to a thirty-day refresh has priced in an asset that depreciates faster than any security program guarding it can mature.
♦
In 1955 Soviet bombers flew repeated circuits past the reviewing stand at the Moscow Aviation Day display, and American assessments of Soviet long-range aviation inflated accordingly. The bomber gap drove procurement. Within two years came the missile gap, which drove more. Both estimates emerged from a system in which the Air Force had institutional interest in the number, contractors had commercial interest in it, and the adversary had operational interest in inflating it. They were corrected by CORONA photography. A new way of seeing did the work that a decade of reasoning had not, and it arrived after the procurement money was already spent.
Twenty years later, Team B was constituted specifically to produce a higher threat estimate than the intelligence community’s own. Its projections were later shown to be systematically wrong; it reset American defense policy anyway.
The lesson is not that officials lie. Absent independent measurement, the estimate that prevails is the one with the best-funded circulation, and skepticism has never been a match for that. Instrumentation has. So the question nobody is pressing with any seriousness is what the CORONA of this would be, the instrument that lets a disinterested party test a capability claim instead of taking it on the vendor's word. Compute accounting is the candidate with the right physics, which is why export controls are the only clause in the Sanders bill that would actually matter.
♦
The public standings mislead for a second reason, older than any of the labs. In the intelligence business, the operational art has always been use without disclosure. Advantage frequently consists in never revealing that you have one. Bletchley Park’s product was protected by an elaborate architecture governing what could be acted upon and what had to be left alone, as acting on everything would have burned the source. Should a state today achieve decisive superiority in cryptanalysis, network exploitation or targeting, the expected observable is silence. Treat the public leaderboard as the contest and you will have mistaken the shop window for the vault.
The mobilization is loud. A model pushed to three million personnel cannot also be the thing hiding, but two different assets are in play. Commodity capability is the model on every desk, loud by design, its whole purpose being scale. The other asset is not three million people asking a chatbot to tidy their email. It is a single air-gapped machine reading through a target's unpatched vulnerabilities and producing an exploit no defender has seen before, or an intrusion that adapts faster than the team hunting it can name what they are looking at. That kind of advantage stays hidden because using it loudly would reveal it and let the target shut the door. None of it appears on a leaderboard. It appears, if at all, as an unexplained breach months later, or as nothing. The benchmark wars, the rankings, the whole visible contest belong to the commodity layer. The advantage that would decide something belongs to the quiet one, and it will never post a score.
From the early 1950s until the mid-1970s, the Soviet Union directed microwave energy at the United States Embassy in Moscow. American research into the effects proceeded under classification for years and embassy staff went uninformed. Both governments stayed quiet. Half a century on, the government that studies the injuries now called Havana Syndrome, funds the research, renamed a directed-energy team after them and pays the victims, still declines to say who or what was responsible. Whatever the cause, the posture is the tell.
The point generalizes past any single case. A capability that confers true advantage buys the silence of the government that wields it, so the loudest signals in a contest are the ones with the least riding on them. Apply that to the AI leaderboard and the ranking inverts; the models competing in public are the ones cheap enough to show off.
♦
If true operational advantage hides in the dark, the visible competition has relocated wholesale to the infrastructure layer, where it is not only loud but staged to be seen. On that prestige competition, the race frame gets the sociology right and the motive wrong.
CNAS now tracks 184 government-backed sovereign AI projects across 67 countries, up from a single project in January 2023. This year the mix shifted: the share of new projects focused on infrastructure rather than models rose to 80 percent in the first half, from 62 percent in the preceding six months. Governments have stopped chasing a national champion model for the flag value and started buying data centers.
That resembles an attempt at dependency management — yet Nvidia supplies hardware for 45 percent of tracked sovereign projects. Mistral, Europe’s designated champion, runs its flagship facility south of Paris on 13,800 Nvidia Grace Blackwell processors with another 18,000 being procured on its behalf, funded partly by a European public vehicle and led by a Samsung investment. India has committed $1.25 billion to the IndiaAI Mission and open-sourced competitive Indian-language models, all of it on American silicon.
Sovereign AI is prestige competition that deepens the dependency it advertises escaping. Reach beyond Sputnik for the precedent and you land on Huawei and 5G, a contest over who supplies the substrate on which everyone else’s sovereignty gets built, where the winner sets the defaults for a generation without ever having to win an argument.
♦
None of this establishes that the competition is fake. Chinese open-weight models now trail the closed frontier by months rather than years, and a safety nonprofit reported in August that one model refused none of the offensive cyber and biology tasks it was given. The gap between the frontier and a competent fast-follower is real, and the historical record shows the odds of successfully protecting an asset one person can carry out in their pocket are not good — but the gap numbers wreck the smuggling story. If the frontier leads by four months and open-weight releases track that close behind, the recruited insider is stealing something the incumbent was going to publish by summer. The heroic security program guards a secret with the shelf life of milk.
Forget the spy. The model release calendar is the channel that moves capability. The most consequential transfer of the last two years was DeepSeek shipping weights that anyone could download. No budget on earth defends against a competitor’s decision to give its work away. Theft is a real problem sitting on top of a larger one nobody files as a security threat at all because the party doing the proliferating is the incumbent, on purpose, as a business strategy.
Old school security playbooks were built for a world of scarcity. Their entire premise is that the dangerous thing is rare and lives somewhere specific, so the job is to keep the adversary out of the room where it is kept. Today, open weights dissolve that premise. When a capable model is posted to a public repository and mirrored across machines, there is no room and nothing to break into. The adversary does not exfiltrate the weapon. They download it, legally, alongside every graduate student and every criminal group on earth. Capability that a year ago would have been a bespoke state program has become a baseline anyone can stand on. A tool that drafts flawless phishing lures in a language the sender does not speak, or generates malicious code on request, has become ambient weather. The security question is no longer how to keep the weapon out of enemy hands; every hand already holds it. The question now is how to survive a world of cheap, endless, automated offense.
♦
The claim that we cannot regulate or we lose to China has exactly the grammar of the claim that we could not halt testing or we would lose to the Soviets. That does not make it false, but does mean it arrives pre-loaded with interest, and the analyst’s obligations follow mechanically: ask who funds its circulation, ask what regulatory action it is timed against, ask whether the adversary capability estimates supporting it originate with parties who profit from them.
The obligation runs in every direction, including those that do not flatter. Safety institutes, academic risk researchers, journalists who cover this beat and an intelligence community seeking appropriations all have an interest in the salience of the threat. The labs distort in two directions at once, inflating capability for the investors and playing down risk for the regulators, and both moves serve the same balance sheet. That regularity is usable. Discount the capability claims, discount the safety assurances, and pay closest attention when a single company is making both on the same day.
Everyone in this argument is describing AI as something akin to a nuclear warhead. There is no warhead. There are files that copy easily, a power and water bill the size of a small country’s, a Barrier Removal Board, and a set of difficult-to-verify capability estimates. The hard problem was never how fast the technology moves. It is that we are asked to believe an account of a race none of us can see, assembled by the only people positioned to run it, at the exact moment the account is being used to clear away the rules.
A memo that writes “AI arms race” into a procurement clause and deletes the safety terms in the same breath is that entire maneuver compressed into six pages and a signature. Naming the race neatly sidesteps the rules, and that maneuver is not a question about artificial intelligence.
It is about the limits of knowledge under adversarial and commercial incentives.
♦




The race is not about the limits of knowledge, rather it is about the impact of consequences. It is the bragging rights of a macho culture that we should be fearing.