JOVANA
Explore Library Glossary Getting Started Three Levels Fields How it works Mission
Join the mission
All guides

International Coordination & What You Can Do

Even a lab that does everything right cannot make the world safe on its own — because it is not the only lab. This capstone is about the layer above any single model: racing dynamics, what coordination really means, the real attempts of the last two years, and the honest, concrete ways an ordinary person can still matter.

The problem no single lab can solve

By the time you reach this guide you have seen, one layer at a time, how a lab can try to make its own model safe: shape its behavior with RLHF, read its internals, test it with evals and red teams, and write the responsible scaling policy that says what it will and will not deploy. Here is the uncomfortable fact that sits on top of all of it. Even a lab that does every one of those things perfectly cannot, by itself, make the world safe — because it is not the only lab, and the others answer to no one's safety policy but their own. This last guide is about that layer: the part of the problem that is not technical at all, and the part where a person who will never train a frontier model — quite possibly you — can still matter.

Picture three teams in a tunnel, racing to the far end where a prize waits. Each driver would rather slow down at the blind curves to check the road — but slowing down means losing, and losing means the prize, and the steering wheel, go to whoever was willing to take the curve blind. So everyone takes the curve blind, not because any driver is reckless by nature, but because the structure of the race punishes caution. That is the shape of the AI safety problem at the level of the whole world. No amount of careful driving by one team fixes the shape of the race itself, and the shape is what this guide is about.

This is why the field has a whole vocabulary that lives above any single model — AI governance, compute governance, and the thing we will spend this guide on, international coordination: getting different countries and labs to move together rather than at cross-purposes. We will work through why a race makes everyone less safe, what coordination actually means in practice, the real attempts of the last two years, the historical analogies people reason from, and finally what an individual can do. Throughout, hold the field's honest uncertainty: whether this whole layer can be made to work at all is one of the most genuinely contested questions in AI safety, and we will not pretend otherwise.

Why a race makes everyone less safe

Define the thing the tunnel was a picture of. Racing dynamics is what happens when several actors compete to build and deploy powerful AI fast, and the competition itself pushes each of them to spend less on safety than they would if they were alone. The logic is a collective-action trap. Safety has a cost — call it a safety tax: the extra months of testing, the capabilities you decline to ship, the outside evaluations you wait for. If a competitor skips that tax, they ship sooner, capture the market and the funding, and set the norms. So even an actor who sincerely cares about safety faces a brutal incentive to cut it — not because they stopped caring, but because caring unilaterally just hands the lead to whoever cares less.

                         OTHER LAB
                   invest safety     skip safety
                +----------------+----------------+
  invest safety | both safe,     | you slow,      |
   (this lab)   | both slower    | they ship 1st  |
                | = GOOD         | = you lose     |
                +----------------+----------------+
  skip safety   | you ship 1st,  | both fast,     |
   (this lab)   | they lose      | both risky     |
                |                | = BAD          |
                +----------------+----------------+

Whatever the other lab does, each lab's tempting move is "skip".
So with no coordination both skip -- and everyone is worse off
than if they had both invested. That gap is what coordination
tries to close.
A stylized prisoner's-dilemma sketch of the race. It is a cartoon, not a claim about any real company: it shows the structure (each side is individually tempted to defect, so both end up worse) that coordination exists to fix.

Two honest caveats keep this from hardening into a slogan. First, the matrix is a cartoon: real labs are not identical players in a one-shot game; reputation, liability, and regulation all change the payoffs, and a lab that has built a serious safety brand may find safety profitable rather than purely taxing. Second — and this matters — not everyone agrees the race is the dominant force. Some argue that a clear, responsible leader with a comfortable lead is safer than a tight pack, because a leader can afford the safety tax; on that view, slowing the leader to let others catch up could make things worse. Whether competition is mostly a race to the bottom or sometimes a race to the top is itself contested. Hold the trap as a serious risk, not a proven law.

A balance note before we go further. Racing dynamics is a structural argument, not an accusation, and it cuts in more than one direction. The same competition that can erode safety also drives the investment that funds safety teams, evals, and interpretability research. The educated reading is not "competition is evil" but "competition shapes incentives, and good outcomes need those incentives steered" — which is exactly the job coordination is trying to do.

What "coordination" actually means

The word "coordination" frightens people because they hear "one world government." It means nothing of the sort. Coordination is a spectrum, a ladder from very weak and easy at the bottom to very strong and hard at the top, and almost all of the real action lives on the lower rungs. The point of the ladder is leverage at acceptable cost: the higher you climb, the more it can actually bind behavior, but the harder it is to get everyone to agree and the more it costs to enforce. Most of what exists today is near the bottom.

  1. Shared language and information. Common definitions of "frontier model," shared incident reporting, published model cards and safety frameworks. Weak on its own, but nothing higher works without it.
  2. Voluntary commitments. Labs publicly promise to do (or not do) specific things — the Seoul Frontier AI Safety Commitments are the clearest example. Binding only on reputation, but reputation is not nothing.
  3. Standards and evaluation. Common safety evals, shared test methods, and independent evaluators who can check a lab's claims instead of taking its word.
  4. National regulation. Laws that bind within one jurisdiction — the EU AI Act, executive orders, national reporting rules. Real teeth, but only inside one set of borders.
  5. International agreements. From non-binding declarations of shared concern up to, in principle, treaties with verification and enforcement. The top rung exists for nuclear weapons; for AI, nothing binding-and-verified exists yet.

If you want to know why one rung of that ladder gets so much attention, look at the enforcement problem. Treaties bite only if you can verify them, and software is almost impossible to verify — weights are just numbers, copyable and hideable, and an algorithm is an idea. But the hardware those models train on is none of those things. The most advanced AI chips come from an astonishingly concentrated supply chain (a handful of firms in chip design, fabrication, and the machines that make the machines), a large training run is physically big, power-hungry, and expensive, and chips are objects you can count, locate, and control. That is why compute governance is the one place coordination already has real leverage: export controls on advanced chips are a live, working instrument, and "on-chip" verification and compute accounting are active research precisely because compute is the rare part of the system you can actually meter.

The other backbone is institutional. Starting with the UK in late 2023, several governments built AI safety institutes — public bodies with the technical staff to test frontier models themselves, rather than taking a lab's marketing at face value. A national AI safety institute can run dangerous-capability evals, red-team a model before release, and feed what it finds back into regulation. Wiring these institutes into a network is an attempt to make evaluation portable across borders — so that a test done in one country counts for something in another, which is the quiet machinery any future agreement would have to run on.

A worked example: the last two years of real coordination

This is not hypothetical; a striking amount happened between 2023 and 2025. In November 2023 the United Kingdom hosted the first global AI Safety Summit at Bletchley Park — the codebreaking site from the Second World War — and produced the Bletchley Declaration, signed by around twenty-eight countries plus the European Union. Its significance was less in what it committed anyone to do (very little, concretely) than in who agreed: both the United States and China put their names to a statement that frontier AI carries potentially catastrophic risks that demand international attention. Getting those two governments to sign the same page about anything is the kind of thing that, a few years earlier, sounded naive.

Six months later the follow-up summit in Seoul produced the Frontier AI Safety Commitments: sixteen companies — including major labs from the US, UK, China, and the UAE — voluntarily committed to publish safety frameworks defining risk thresholds and what they would do if a model crossed them. If that sounds familiar, it should: it is essentially the responsible scaling policy idea from the previous guide, with its if-then commitments, lifted from a single company onto a semi-international stage. A growing set of countries is also assembling a network of the AI safety institutes described above, so that evaluations and methods can be shared rather than each nation rebuilding them from scratch.

Two more pieces complete the picture. The International AI Safety Report, chaired by the deep-learning pioneer Yoshua Bengio and backed by around thirty countries plus bodies like the UN, EU, and OECD, published its first full edition in 2025. It is a deliberate echo of the IPCC for climate: not a policy, but a shared scientific baseline, so that argument can proceed from common evidence rather than from everyone's favorite anecdote. Alongside the summits sit instruments with actual force: the EU AI Act (adopted in 2024, risk-tiered, phasing in over years) is binding law within Europe, and US export controls on advanced chips are compute governance operating in the real world right now.

Historical analogies — and where each one breaks

We have coordinated on dangerous technology before, and those precedents are how serious people reason about whether AI coordination can work. Each analogy is instructive — and each is instructive precisely where it fails. The discipline is to take the lesson and notice the break in the same breath, because a half-remembered analogy is how confident, wrong conclusions get made.

The cheerful one is the Montreal Protocol of 1987, which phased out the chemicals destroying the ozone layer and is the great success story of environmental coordination. Why did it work? The harm was narrow and measurable, only a handful of chemicals were involved, viable substitutes existed, and the cost-benefit case was clear. AI breaks the analogy on every count: the "harm" is diffuse and disputed rather than a single measurable hole in the sky, the technology is general-purpose and economically irresistible rather than a niche refrigerant, and there is no clean substitute to switch to. Nuclear arms control is the harder, more apt precedent — and it teaches the verification lesson. The NPT, the IAEA, and test-ban treaties have real teeth because enrichment leaves physical, detectable traces: inspectors can verify. AI's nearest analogue to fissile material is compute, which is exactly why compute is the live lever — but weights and algorithms are software, copyable and concealable in ways uranium will never be.

Two cautionary precedents round it out. The Biological Weapons Convention of 1972 banned an entire class of weapons but was written with essentially no verification regime, and it is widely judged weak for exactly that reason — a treaty you cannot check is closer to a promise than a constraint. And the 1975 Asilomar conference, where biologists voluntarily paused and self-regulated the frightening new technology of recombinant DNA, is the optimist's favorite: proof that a research community can govern itself early and responsibly (the 2017 "Asilomar AI Principles" deliberately borrowed the name). The thread running through all four is blunt and useful: coordination has genuinely worked when harms were measurable, verification was physically possible, and substitutes existed — and AI is hard on all three counts at once. That is the honest reason thoughtful people land in such different places on whether it can be made to work.

Common misconceptions & pitfalls

The first misconception is that coordination means a single world government deciding what everyone may build. It does not; almost everything real lives on the lower rungs — shared definitions, voluntary commitments, common evals, verification — and no serious proposal asks for, or needs, a planetary sovereign. The second is the fatalist's line: "China will never cooperate, so it is pointless." That is too neat. China signed the Bletchley Declaration, has its own AI regulations (stricter than the West's in some areas, such as algorithm registration), and shares the catastrophic-risk language in official statements. Cooperation with rivals is hard and partial, not impossible — and the mirror-image error, assuming cooperation is easy or already secured, is just as wrong.

The third misconception is that regulation means stopping AI. Almost every serious proposal targets a tiny number of frontier models — the largest, most capable training runs — and leaves the vast ordinary use of AI untouched; the frontier AI regulation debate is about the handful of biggest systems, not your email autocomplete. The fourth is mistaking a summit or a declaration for a solved problem: those are shared concern, not enforcement, and reading a non-binding statement as a verified, binding treaty is the most common over-optimistic error in this whole area. The fifth, and the one that matters most for you, is believing an individual cannot possibly affect any of this.

What you can do

You do not have to train a model to matter. AI safety is a young field with far more open problems than people, and most of its leverage is not in writing loss functions — it is in governance, security, evaluation, communication, operations, and plain clear thinking. Be honest with yourself about the uncertainty, though: nobody can promise that any one person's effort changes the outcome, and the field is full of disagreement about which efforts even help. What follows is not a recipe for guaranteed impact; it is a way to put your weight where it has the best chance of counting.

  1. Calibrate your beliefs first. The single most useful contribution is to think clearly: hold demonstrated facts tightly and speculative conclusions loosely, and resist both breathless hype and reflexive dismissal. You already practiced this across the whole ladder — it is itself a contribution to a debate poisoned by both extremes.
  2. Skill up where you already lean. ML and research for technical alignment and interpretability; policy, law, and economics for governance; security for evals and robustness; writing and teaching for field-building and clear public communication; operations and management to make safety organizations actually function. The field needs all of these, not just researchers.
  3. Pick one lever, not all of them. Technical safety, governance, security, field-building, communication, or a supporting role — depth in one beats a thin layer spread across everything. Most people who shaped this field went deep on a single thing.
  4. Act where you already are. Inside a company: push for evals, red-teaming, and honest disclosure. As a citizen: cast informed votes, comment on proposed rules, support competent institutions. As a builder: deploy your own systems carefully — you learned exactly how across this ladder, from RLHF to red-teaming to safety cases.
  5. Stay honest, and stay in it. The two failure modes are burning out on doom and drifting into cheerleading; the contribution that compounds is the patient, calibrated kind that is still there in five years. Endurance with good epistemics beats a brief, loud sprint.

Most readers will contribute indirectly, and that is genuinely valuable: by being a clear-thinking voice in a noisy room, by raising the safety baseline wherever they happen to build, by funding or supporting good work, by simply not adding to the hype in either direction. The differential technological development idea from the foundations rung applies to a career as much as to a research portfolio: where you have a choice, steer your effort toward the understanding-and-safety side of the ledger relative to raw capability. You will not always have that choice cleanly — but noticing when you do is most of the work.

What's still debated, and where to go next

The first open question is whether meaningful coordination is even feasible. Pessimists argue that the race plus great-power geopolitics make deep, verified cooperation a fantasy — expect at best voluntary theater that evaporates the moment it costs someone the lead. Optimists point to ozone and nuclear as proof it can be done when stakes are clear enough, and note that the last two years moved faster than almost anyone predicted in 2022. Both are reasoning honestly from real evidence; the question is genuinely unresolved, and your own read on it should stay tentative.

The second debate is sharper: slow down, or speed up? In 2023 an open letter called for a six-month pause on training systems more powerful than GPT-4; thousands signed, and many equally serious researchers refused, calling it unworkable or even counterproductive. Slogans have grown up around the poles — "d/acc" (a differential, defense-favoring acceleration) and "e/acc" (effective accelerationism, roughly: more powerful AI sooner is net good). Underneath the noise sits the real driver: people's timelines and their p(doom) estimates differ by orders of magnitude, so from the very same facts they rationally reach opposite policies. There is no neutral ground here that everyone would accept, and pretending otherwise is its own mistake.

The third debate is whether regulation is cure or capture. Critics warn that safety rules written around frontier models entrench the incumbents who can afford to comply (regulatory capture) and crush open-source and small players; defenders answer that the largest training runs are exactly where oversight belongs, and that fully open weights would remove the compute lever entirely. Cutting across all of it is the framing war: the national-security frame — "we must win the race against rival X" — directly opposes the cooperative frame, and which one dominates the conversation may matter more than any single law. These are disagreements about both values and facts, not puzzles with a settled answer, and you should expect thoughtful people to remain split.

This is the top of the governance rung and the end of the ladder — but not the end of the field. Three directions from here. Loop back: reread the foundations rung now that you have seen the whole arc; the case for catastrophic risk and even the alignment problem read differently from the top. Go deep: pick the track that gripped you most — interpretability, evals, RLHF, governance — and follow its primary sources past these introductions. Stay engaged: read the International AI Safety Report, follow what the safety institutes publish, and keep your epistemics honest in public. You arrived an absolute beginner; you leave research-literate. The last and most durable thing you can do is keep thinking clearly, out loud, in a field that badly needs exactly that.