What, exactly, would we have to stop?
In one of our more important write-ups early this year, we discussed how AI is going to be our economics and AI is going to be our politics for years and decades ahead. The debate over controlling it is making that plain. We normally prefer to reach our conclusion with the reader. With so much being written, we are making an exception: a lasting global ceiling on AI’s capabilities would require control reaching across the whole system that produces them. Models and their use, but also processors, accelerators, memory, interconnects, and the software that organizes them, if not even the humans who are allowed to use them. Restrictions confined to models or today’s arrangements for using them cannot guarantee that ceiling while the rest keeps advancing.
Dario Amodei writes in We Must Pace the Frontier: “We must slow the pace at which we improve the capabilities of AI models.” His proposal includes external evaluators working inside laboratories, safety requirements tied to capabilities, and international coordination. He also considers training inputs and advocates chip export controls. He is proposing measures to give safety work time to catch up, and acknowledges how difficult comprehensive global restraint would be.
Others have more points to add. Many of these are substantive proposals. They deserve to be discussed on their actual terms. The range is wide. In 2023, Eliezer Yudkowsky called for an indefinite worldwide moratorium on large AI training runs. Other approaches concentrate on particular applications, users, or harmful activities. Between them lie different combinations of development restrictions, deployment conditions, hardware controls on people the proposers do not trust, and ongoing supervision. Agreement that AI needs control does not settle which of these ambitions people mean.
The shrillest cries are coming from the late converts: those who were saying AI is all hype and isn't being used by anyone are using the debate to maintain their negative view of innovation industries’ prospects by raising the specter of controls. This article is not about that. The issues being raised are far more serious than converting them into an argument about stocks.
It is good that capable people are working on these questions. Much of that work already goes into considerable detail. A measure intended to buy time should be judged on whether it buys time. Our question concerns the broader challenge of containing AI’s expanding capabilities worldwide.
Here is the test we propose: everyone agrees, everyone complies, and the models stop improving. What can still happen?
We will move from agreement to legal language, from language to measurement, and from measurement to enforcement. At each stage, the same difficulty returns: capability can grow through changes outside the particular activity being restricted. Researchers recognize these effects. We want to follow their implications through to the agreement that would be required.
Our examples are deliberately simple and their numbers hypothetical. Real systems have dependencies and bottlenecks. The examples isolate a question that complexity can otherwise conceal: how much can change while the model stays the same?
Agreement is the easy sentence
“AI must be controlled” attracts broad support partly because everyone can supply their own meaning. So does “crime must be reduced”. Neither sentence tells a police officer what to do tomorrow morning.
During COVID, a shared desire to prevent deaths did not settle when to close schools, restrict travel, or reopen workplaces. There were genuine uncertainties and competing costs. Policies diverged even within countries and counties. They changed with changes in politicians. Agreement on the desired outcome did not settle the permissible means.
Climate negotiations have produced meaningful agreements. Yet a common global goal still leaves countries deciding how to divide costs and change their economies after decades of arguments. The Paris Agreement combines international commitments and review with nationally determined plans. Reaching an agreement and implementing a uniform restriction are different achievements.
AI adds a difficult conflict. What one government wants slowed may be what another sees as its chance to catch up. Restricting another country’s access signals distrust and gives those excluded a stronger incentive to develop their own alternatives. Within countries, a firm stance by one political party gives its opponents an opening to champion a different course. Companies will find opportunities in the uncertainty, pressing for rules that constrain their competitors more than themselves. Those disadvantaged will question whose interests the restrictions serve, with accusations of bias and discrimination. The AI control debate is sure to lose its primary purpose soon in our world of competitive nations, corporations, and political landscapes.
Politicians answer to their own electorates. Companies answer to customers, owners, and employees. The benefits surrendered may be immediate and local; the dangers avoided may be uncertain and shared. The same person can sincerely fear uncontrolled AI and sincerely fear what falling behind would mean for their country.
Now extend the proposed controls to memory, processors, and communications (this is separately discussed in the sections below about how the controls cannot be simply about the model developments). More industries and competing needs enter the negotiation. Equipment that makes agents more effective may also improve medical imaging or reduce electricity consumption. The scope of the bargain grows with the scope of the promised control.
These are formidable problems. For the rest of the argument, let us assume we have solved them.
Now write the rule
Everyone has agreed. A lawyer must now turn that agreement into a restriction: “Do not develop AI beyond the agreed frontier.” Eventually, someone may stand in a courtroom accused of crossing it. What, precisely, would have to be proved?
Ever since the arrival of ChatGPT, following model rankings can feel like following the weekly rankings of the music hits, with several charts announcing different number ones. The broad advance is clear. Which model leads, by how much, and in what respects is less straightforward. Different benchmarks can tell different stories, particularly when comparing closely matched models.
Imagine one model leads in mathematics while another is better at completing a day’s computer work without help. Which is more capable? Combining their results requires deciding how much each ability matters. Choosing which results justify a restriction requires another judgment. The score can be precise while the choices behind it remain subjective.
“Capable of” is an unfinished sentence. Doing what, with which tools, over how long, and with what reliability? Give the same model more attempts or a better way to check its work, and its measured performance may change. Those conditions become part of what we are measuring, and therefore part of what the rule must address.
Defining “the frontier” is harder still. Does crossing it mean surpassing any existing ability, achieving a higher average score, or acquiring a particular dangerous capability? These describe different boundaries. A system could cross one while remaining comfortably below another. The word “frontier” does not tell us which boundary matters.
Even the measuring instruments keep changing. Older tests become too easy for leading models to distinguish; harder ones take their place. In June, Artificial Analysis removed a benchmark from its main index because it no longer sufficiently separated frontier models, while revising other tests and their weights. Maintaining a useful measure requires changing it.
Now attach legal consequences to such an index. Change the tests or their weights, and an unchanged model could fall on a different side of the restriction. Keep them fixed, and the rule risks becoming obsolete. Who may change the measure, on what evidence, and what happens to an approval already granted? These choices determine who may continue working and who must stop. Imagine writing laws for a system whose measuring instruments change and must be agreed upon by any assigned body every few weeks.
Serious researchers and legal scholars are working on these questions. Better assessments can narrow the uncertainty. They cannot remove the need to decide which abilities matter, how much evidence is sufficient, and whose judgment prevails when the answers differ. Everyone may have agreed to control AI. The language of that agreement still has to survive a dispute between people whose understanding of it is constantly evolving, likely at a pace far slower than changes in capabilities.
Purpose and responsibility require further choices. Software analysis can help someone repair a weakness or exploit it. A rule against unauthorized intrusion addresses a recognizable act. Restricting a capability that might assist intrusion reaches further into legitimate work. The law must also say what is expected of the model developer, the operator, and the people supplying its tools.
We have language for particular activities and methods for particular assessments. What we lack is a settled, general definition of dangerous capability that remains dependable across changing combinations of models, equipment and use. Adding “under foreseeable conditions” creates another question: foreseeable to whom, with what knowledge, and at what point?
Every important word left undefined is worse than a decision postponed. They will simply provide the smarter operators to skirt around them. A restriction that leaves everyone free to decide what it restricts can attract widespread agreement without requiring much restraint. Until we specify what someone must actually stop doing, we cannot know how much we have agreed to control.
Freeze the model. Continue everything else.
Imagine every model is frozen today. No retraining. No new generation. The weights remain identical. Everyone complies.
Consider the scale of the recent Navier-Stokes effort. OpenAI reports that roughly 10,000 cooperating agents reached a proposed solution to the Millennium Prize problem in 88 hours, followed by another 17 hours of formal verification. Millions of messages passed between agents as they explored approaches and shared findings. The model was also upgraded during the effort, so this result does not isolate the contribution of hardware. It does, however, give us a concrete picture of how much investigation can now fit into a few days.
Now freeze that final model. For our thought experiment, imagine running it on a slower, more limited installation a year earlier. Fewer approaches can be explored simultaneously. Each round of investigation takes longer. A promising argument may remain unfinished when the available time or money runs out. We cannot claim that this is what would have happened to Navier-Stokes. But an unfinished investigation would hardly prove that the model could never reach the answer.
Then let hardware advance again. The same weights can support a broader search, more checking, and longer chains of work within those same four days. An approach that previously demanded an impractical commitment may become worth attempting. The difference could be whether an investigation produces another inconclusive report or an entirely new result.
Harmful uses that lie beyond practical reach today may therefore become feasible with the same models on better hardware. A freeze on model development can be fully respected, even as the capacity to cause harm continues to grow.
A lasting ceiling on AI’s capabilities would therefore require controls reaching beyond model development into the hardware that expands what those models can do. Processors, memory, and interconnects all enter the argument. Every difficulty discussed above, from agreeing on what to control to defining and measuring it, would have to be confronted across these fields too. Containing AI becomes a negotiation over progress across entire industries, each with purposes, customers, and national interests extending far beyond AI.
This is the difficulty of control. We cannot infer from what a model failed to achieve on yesterday’s equipment everything it might achieve on tomorrow’s. Nor can we prepare a complete list of the discoveries that longer, broader investigations will uncover.
A limit on what?
Suppose our rule permits no more than 100 agents to work together. What counts as one? A program, a task, a separate working history? The number becomes meaningful only after we settle what is being counted.
Make the rule more precise: count simultaneous model requests. Faster hardware can finish each batch sooner while respecting the limit. Cap total computation, and better software can achieve more within the allowance. Each restriction constrains something measurable. Neither necessarily holds useful capability still.
Developers could also use AI itself to find these opportunities: reorganizing work, reducing the computation a task requires, and discovering combinations the rule did not anticipate. A precise limit can become an engineering target, with considerable ingenuity devoted to accomplishing more beneath it.
Benchmark restrictions introduce a further incentive. Today, a high score helps sell a model. Training can be tailored to particular tests, improving the score without a comparable improvement in broader ability. Researchers have demonstrated how pronounced that distortion can become.
Now make a high score the trigger for restrictions. The commercial incentive can reverse. A developer may want the model to become more capable in practice while appearing less capable in the assessment. The ambition would be to improve what customers can accomplish without raising the number that attracts regulatory attention.
This possibility has already been explored experimentally. Researchers have trained models to conceal particular capabilities during evaluations while retaining them under other conditions. They have also demonstrated that models can be prompted to target particular evaluation scores. These are controlled demonstrations, but they establish that measured weakness need not mean a lack of ability.
A regulatory framework would therefore have to distinguish genuine limitations from lawful improvements within the rules and deliberate attempts to mislead assessors. Independent testing can make concealment harder. Writing a precise threshold does not, by itself, resolve any of these problems.
Beyond individual model makers
A provider may monitor every request and still see only part of the work. Someone intent on harm need not keep the whole project inside one model, one company, or one supervised environment.
This already has a concrete example. Anthropic’s September threat report describes a deceptive dating-app operation that assigned different roles to different AI providers. Claude supplied conversations, another model assisted human workers, and an image model supplied avatars. Anthropic banned accounts and worked with other laboratories to disrupt the operation. Several systems contributed different capabilities to the same deception.
Now consider a thought experiment involving a fraudulent investment business. One commercial model helps analyze public company reports. Another translates ordinary correspondence. A third helps maintain the website. The operator brings these contributions together privately and uses them to support a dishonest offer. The research, translation, and programming requests need not individually reveal the fraud they serve.
These are simple examples. The broader point is that a mixture of models is capable of creating a combined impact that is materially larger in scale. And this capability can go more undetected with the intermingling of open-source models run on private machines.
This last point needs more elaboration. AI work can take place beyond the monitors of commercial providers altogether. A downloaded model can run on a personal computer or private server without reporting its activity to its developer. Local operation can be entirely offline.
The object of control has expanded again. It includes a project assembled across models, services and private equipment, potentially spanning several jurisdictions. Its most consequential feature may be the relationship between its parts. That relationship may be visible only to the person putting them together.
The freeze that keeps growing
There is a final difficulty. We do not know everything that today’s intelligence can eventually be organized to do. A useful method may be discovered through experimentation long after the model was released. A previously harmless deployment may gain a consequential capability through a new combination of familiar parts.
That creates a continuing regulatory task. Which discoveries require disclosure? Which upgrades require reassessment? How much freedom does an approved user retain to improve the work? A tightly specified deployment can be supervised, but the approval of that deployment cannot settle the status of every later arrangement built around it.
This also matters economically. Even with model development frozen, more people could adopt AI, existing users could assign it more work, and bad actors could attempt projects previously too expensive or too unforeseeable to consider with the availability of information on capabilities. One person providing a solution to a seemingly harmless problem in one part of the world could be the missing key for a player intending to harm.
Greater adoption and more intensive use could increase capabilities beyond those provided by model makers. A pause in model progress does not settle a pause in rising abilities simply through greater adoption. At the outset, the controllers have no way to know how the abilities could keep exploding simply through more adoption.
The argument that will stay with us
One can go on and on with the issues surrounding AI regulations and control. However, our reservations should not mask the main point. We agree that AI needs regulation. We expect the subject to remain with us for years and decades. The range of positions will keep shifting as harms become clearer, safeguards are tested, and, along with AI’s useful applications, its harmful impacts become familiar, as we highlighted in our year-old articles on AI’s virtues and vices (different articles).
Many of the above seemingly difficult issues have solutions, and they will emerge as debates move beyond statements of intentions for grandstanding. Our own views will surely change as we learn.
Still, it is important to also recognize that AI control will be a constant topic, with perpetually changing details, in the era ahead. Every substantial rule distributes benefits, costs, and opportunities. Extremely few will be accepted universally. And very few will remain intact and practically useful for long.
Making it more difficult, the topic will also attract players with different motivations and vested interests. For an increasing number of politicians, AI control slogans will be the most defining feature of campaigns. For corporates, engineering the right laws will be a strategic necessity. For commentators, podcasters, and even late-night TV hosts, coming up with the wildest ideas could provide the best fodder for virality.
One concern will be particularly persistent. The organizations most willing to explain their work will often be the easiest to constrain. Those beyond effective scrutiny may remain harder to reach. Essentially, controls may contain the good guys far more, particularly in their ability to provide the best defenses against those intending to harm. People will ask whether restraint is making them safer while less-accountable competitors grow stronger. That question can arise from a sincere concern for safety as well as commercial self-interest. In a world of little trust, and hence little universal cooperation, unilateral impositions by the stronger players will increase the resolve of those behind to develop more aggressively outside the fold.
That said, innovation, not just in model making, will have to adapt to these changing conditions. Containing AI’s overall capabilities reaches into hardware, memory, interconnects, tools, and working methods, across industries and borders. We expect the conversations to spread beyond the pace-slowing arguments at the model-making level.
Calls for control are now an industry of their own. Serious research and practical safeguards must be distinguished from the volume of statements surrounding them. Agreement that something must be done is the beginning of that work; it is unlikely anyone disagrees, regardless of the tone of those braying for it, behaving as if they are the only ones. People asking for safety deserve an honest account of what the promise requires. The task of effective control is far more difficult than building the world’s biggest innovations. And in all this, we have not even said a word about the time taken before anything is agreed versus the time we have before the models improve multifold.




