The Frontier Asks to Slow Down, Who Holds the Brakes?
Frontier AI leaders are asking for safety to catch up with capability, but restraint becomes harder when every company and country fears being the one left behind.
What troubled me about the AI incident I began following some months ago was not simply that a model had behaved strangely, but that each attempt to understand what happened led to another question. I followed that trail through The System Did Exactly What It Believed Success Required, Before the First Action and Sandbox Can’t Hold Us Down, trying to understand what it means when increasingly capable systems pursue an objective beyond the assumptions of the people evaluating them. The boundary between a controlled environment and the real world was itself becoming part of the problem, and it left me with a question I could not dismiss: how much does intelligence depend not only on what a system can do, but on whether it understands where it is, what kind of situation it is dealing with and where its authority is supposed to end? I did not think one incident could tell us everything about where AI was heading, but I also did not expect that question to return so directly through the people building the frontier.
Dario Amodei, CEO of Anthropic, has now published We Must Pace the Frontier, arguing that capability should advance at a rate that gives safety, alignment and oversight a chance to keep up. More importantly for the trail I have been following, he points directly to the OpenAI–Hugging Face incident as one of the developments that made the problem more serious for him. He is not proposing that development stop or that we abandon the extraordinary benefits he believes more capable AI could bring. He is asking for enough time between capability advances to understand and evaluate what has changed, beginning with a commitment to give independent evaluators ongoing access to Anthropic’s systems. Sam Altman and Elon Musk have publicly supported his direction, with Altman also committing to independent evaluators, while Demis Hassabis has supported the broader direction. That matters because concerns which could once be treated as matters of alignment theory, speculative risk or strange behaviour during evaluations are now prompting questions about the pace of development itself. People who compete intensely over that pace are beginning to ask whether the mechanisms for understanding and containing capability can keep up with it.
The earlier incidents help explain why that question cannot be answered simply by pointing out that no catastrophe occurred. Limited damage can still reveal behaviour whose consequences would change with the capabilities of the system producing it, and Anthropic’s disclosures about its own models accessing external systems during cybersecurity testing make this harder to dismiss as something peculiar to one laboratory or model family. For me, the important change is not only what the model can do. Give a system stronger reasoning, better coding ability, more autonomy, more tools, longer operating time and wider access, and you change what a misunderstanding, an unexpected strategy or a poorly specified objective can become. A failure that once produced a strange answer can become a sequence of actions, while a system that misunderstands the boundaries of a test becomes much more consequential if it can also act beyond them. The technical problem therefore concerns intelligence and its reach together: not only whether the system can reason more effectively, but what that reasoning can now affect when its understanding of the situation is wrong.
Now consider how that expanding reach combines with another development in Amodei’s argument: AI systems are increasingly helping to build the next generation of AI systems. If capability begins improving capability, the mechanism producing the advances may itself accelerate, leaving less time to understand one generation before the next arrives. This is why I think pacing becomes a more serious question than whether we are comfortable with the latest model. We are trying to understand systems while both their abilities and the conditions in which they operate are changing, and the tools helping us develop the next systems may be shortening the time available to do that work. The question stops being only how intelligent is the model? and becomes what can this level of intelligence reach, how quickly is that reach expanding, and are our methods for understanding it improving at the same speed? Even a convincing answer would leave something important unresolved, though, because understanding how a system pursues an objective does not tell us whether the person giving it that objective wants something we would welcome.
That brings this discussion back to The Frontier and the Rest of Us, where I was interested in capability travelling out of elite laboratories and highly technical environments towards ordinary people. A warehouse worker may understand a broken process, a mechanic may carry years of undocumented knowledge, or a small business owner may recognise a problem without possessing the technical ability to build a solution. Coding, research, analysis and design becoming easier to reach could allow those people to attempt things that once required specialist teams or years of training, and I still believe that possibility may be enormously important. But the rest of us does not describe one kind of human being. The same intelligence enters a world containing scammers, criminal organisations, hostile states, surveillance operators, propagandists and people looking for ways to exploit others. The promise and the difficulty arrive through the same movement: capability is becoming available to people who previously did not possess it, while their intentions remain their own.
Anthropic’s latest threat-intelligence work describes misuse ranging from cyber operations and surveillance to fraud and weapons-related activity, which gives that difficulty a more concrete form. AI does not need to invent a new kind of criminal to change what an existing actor can attempt; making that actor faster, cheaper, more scalable or technically capable may be enough. This is the harder extension of the argument I made in The Frontier and the Rest of Us: if powerful intelligence lowers barriers for constructive ambition, it may also lower some of the barriers for destructive ambition. A person with deep knowledge of a useful problem and a person with malicious intent can both find that the technical distance between intention and action has become much shorter. Democratising intelligence does not democratise intention, and greater access can therefore expand human potential without making every consequence of that expansion desirable.
I do not think we can resolve that tension by deciding access itself is the problem. I would be deeply uncomfortable with a future in which powerful intelligence remains permanently concentrated among governments, corporations and wealthy technical elites while everyone else receives a deliberately weakened version. The possibility of ordinary people doing things that previous technological structures kept beyond their reach does not become less important because malicious people exist, but neither does that possibility remove the harder consequences of wider access. We have to ask whether increasingly powerful capability can be distributed without every dangerous capability being distributed in exactly the same form, where those boundaries should be drawn and whether the safeguards maintaining them can remain effective as models improve. Once we ask who should make those decisions, however, we have to consider more than the behaviour of a model or the intention of an individual. Somebody must be able to impose a limit, and that limit must still hold when the people subject to it have something to lose.
This is where the problem becomes systemic. Suppose the companies agree that a particular advance needs more time for evaluation, or that a capability should not yet be widely released. Who holds the brakes? Anthropic, OpenAI, xAI and Google DeepMind can all support stronger safeguards while still competing for users, researchers, investment, compute, contracts and technological leadership. A delay can affect market position, funding and infrastructure commitments, while governments look at the same capabilities through economic advantage and national security. Those pressures influence the decision to move more carefully. A company that waits may watch a competitor release, and a country that restrains development may watch a rival pull ahead. Recognising a danger and accepting the cost of responding to it are therefore not the same decision, especially when the cost depends on what somebody else does next.
Amodei’s proposal has to address that difficulty. It begins with evaluators that one company can admit, then extends towards coordination among frontier companies and democratic governments, and eventually towards agreements involving geopolitical competitors. Yet his argument also limits how far democratic countries should slow down if that would allow China to overtake them, while any global arrangement depends on verification credible enough to reduce the fear of a rival secretly racing ahead. He later called China the toughest dilemma and acknowledged that a shared speed limit might not be achievable. I think it matters that both positions form part of the same proposal: he wants restraint and he wants democratic countries to retain strategic advantage. That is not a tension we can dismiss as hypocrisy or solve by choosing the more reassuring half of the argument. It is part of what the proposed restraint would have to survive.
For me, this is where we need to slow down becomes much more difficult than it first sounds, because everyone can agree that slowing down would be safer while nobody wants to be the actor who slows down first and gets left behind. A company can sincerely believe that a capability is dangerous and also believe it would be more dangerous if a competitor developed it first. A government can want international controls while refusing terms that might leave a rival technologically ahead. Neither position necessarily requires dishonesty, and that is precisely what makes the problem systemic: rational decisions made separately can produce a collective outcome that many of the participants say they do not want. Agreement about the danger does not remove the incentive to continue, particularly if each actor believes its own continued progress is a defence against somebody else’s.
The political responses to the support from frontier leaders show that this is already more than a theoretical difficulty. Kirill Dmitriev, a representative of President Vladimir Putin, responded that the genie could not be put back in the bottle, presenting a slowdown as unrealistic once the technology was already moving. President Trump resisted substantially slowing American AI development, allowing that guardrails may be needed while framing the issue around competition with China and saying that “whoever wins AI, wins.” Meanwhile, China’s state-run Global Times attacked Amodei’s proposal as a “Cold War playbook” aimed at containing China, particularly given his calls for tighter controls on advanced chips, distillation and model-weight theft. These responses do not establish one shared political position, and state-backed commentary is not the same thing as a formal international commitment. What they show is how quickly a proposal about safety becomes an argument about who retains power, who gives something up and whether the terms can be trusted.
That is why I do not find the picture of responsible companies facing irresponsible governments especially useful here. The same restriction can look like a necessary safeguard to the actor proposing it and a means of preserving dominance to the competitor being asked to accept it. Amodei’s own combination of pacing and strategic advantage makes that tension part of the proposal, so persuading others that the risk is real would not by itself settle the disagreement. Different actors can recognise parts of the same danger and still disagree sharply about what responsible action requires. The hardest part may not be getting everyone to recognise the danger. It may be getting everyone to accept the cost of restraint when nobody can guarantee that everyone else will pay the same cost.
Once the problem is understood that way, who holds the brakes? becomes a question about authority rather than an appeal for somebody to be more responsible. If frontier companies control them, they are being asked to make safety decisions against their own competitive incentives. If governments control them, national advantage enters the calculation. If an international body controls them, we have to ask who joins, who has authority over whom and what happens when somebody refuses. Answering that question also requires credible verification, because an agreement cannot offer much reassurance if participants suspect a hidden training run, a secret model or a different definition of what slowing down means. The authority to announce a limit is not necessarily the authority to enforce it, and enforcement becomes harder when the capability being restrained may have enormous military and economic value.
This is also where the early agreement among Amodei, Altman, Musk and Hassabis will meet a more demanding test. Their support matters precisely because they are competitors, but safety should catch up with capability leaves open which capability should trigger a delay, how long the delay should last, what evidence permits development to restart and whether everyone accepts the same threshold. Imagine the point at which an independent evaluator says not yet while a delay puts billions of dollars, strategic advantage and months of technical progress at stake. Does that judgement meaningfully delay deployment? What happens if one company accepts it and another decides the evaluator is too cautious, or if a government believes the delay would expose it to a rival? Agreement in principle becomes restraint in practice only when somebody accepts a cost they would rather avoid, and the value of the safeguard depends on what happens at that moment, not only on the sincerity of the original promise.
I do not yet know what arrangement could answer all of those questions, and I do not want to force an answer simply because the people building these systems have begun asking them more urgently. Nor do I think we learn much by either celebrating every warning as proof that the industry has become responsible or dismissing it as an attempt to protect market position. What interests me is whether the system surrounding these actors allows the restraint they are discussing: whether evaluation can genuinely delay an advance, whether competitors will accept comparable limits and whether anybody can establish that the people claiming to have pressed the brakes have actually done so. Those remain open questions, but they now connect much more directly to the earlier incidents than I expected. A question about a system acting beyond the boundaries of a test has led into questions about the reach of intelligence, the intentions it serves and the authority capable of containing it.
That does not settle the argument I have been developing; it makes it more serious. Can technical controls keep up as capability gains greater reach and helps accelerate the next advance? What happens when that capability spreads across every kind of human intention, including the people whose ambitions we would not want to assist? Can the competitive systems building it coordinate restraint when restraint itself carries a cost? These are connected problems without being interchangeable: controlling a model does not settle who should have access to its capabilities, and agreeing where access should end does not establish who can make that boundary hold across a competitive world. If the people running towards the frontier are beginning to ask how to slow themselves down, then the next question is not simply whether they should. It is whether anyone has yet built brakes that the race itself can survive.
— Teff