What does the future look like for international, intergovernmental cooperation in AI? Despite my usual stance as an optimist, I’m skeptical here, particularly when it comes to making the most powerful technologies widely available. I think the ‘dark forest’ theory from science fiction can help explain the challenges.

At the heart of the geopolitical dynamic around AI is tension between the U.S. and China, which blends together issues of openness, security, cooperation, and access to key inputs and resources. Xi Jinping just gave a major speech on the subject at the “2026 World AI Conference and High-Level Meeting on Global AI Governance” in Shanghai, and led by China, 29 countries came together to join the new “World Artificial Intelligence Cooperation Organization”. This piece by Bryan Alexander has a good map of them. Politically, the group feels like an intentional counter to Western (i.e. American and European) interests in AI development and cooperation; I don’t see any future for global legitimacy there.

This effort at cooperation comes on the heels of Chinese AI company Moonshot releasing Kimi K3, a new frontier-class model, for which the company says it will share open weights publicly on July 27. Kimi is, by most technical measures, competitive with (if slightly behind) the best offerings from Anthropic and OpenAI, and far cheaper. To many, “China has all but caught up,” and competing just on the basis of model quality is not a path for the U.S. to “win”. Some commentators suspect a future where open models from China are banned in the U.S. The “Europe 2031” authors envision a future where both China and the U.S. ban the release of open model weights on different grounds.

Are we headed for an AI cold war, as Alexander suggests, with protracted back and forths between two leading world powers? Or is there hope for collaboration? In theory, open weight models are a nonrivalrous good, and shared open source safety tools have massive and global benefits as well (thus the existence of, e.g., ROOST). If we could collectively agree to share open weights and safety tools, and not hog those resources that are rivalrous (chips, rare earth elements, etc.), everyone would be better off.

Certainly to the extent that improvements in AI are motivated by scientific research, there is precedent for intergovernmental collaboration in the form of CERN. It’s a compelling story: the resources required to undertake advanced physics experiments are staggering, and the learnings of them are foundational, perhaps of broad potential use but nothing so clear nor certain to justify risky investment by a single actor. So the best way to make progress was to create and fund an international entity to take the work forward. Similarly, on some level, advancements in AI are staggeringly expensive at the frontier, and some may think the investment/benefit calculus is quite risky, yet improvements can ultimately benefit everyone. Gary Marcus proposed a CERN for AI nine years ago.

As mentioned, I’m a persistent optimist. But I have my limits, and I’m not optimistic here. The baseline is competition; moving from there to cooperation feels difficult under any circumstances, and the world isn’t exactly operating at a particularly collaborative baseline in non-AI contexts right now.

I want to suggest a framework to explain why collaboration in favor of open tools, under any circumstances, is difficult to sustain. In Cixin Liu’s “Three-Body Problem” science fiction trilogy, the reactions of civilizations to discovery of another civilization are cosmo-sociologically governed by what is described as the ‘dark forest theory’. (See, e.g., this independent articulation.) Under the theory, the universe is a dark forest in which it’s difficult to see very far, but after sufficient exploration one might come across another being. Because it is dark and there is next to no information, gauging whether that being is a threat is difficult. And if it’s malevolent, it will strike and harm you before you can defend yourself. Incentives and fear therefore overwhelmingly encourage striking first.

Liu gives two axioms and two concepts for dark forest theory:

  1. Survival is the primary need of civilization.

  2. Civilization continuously grows and expands, but the total matter in the universe remains constant.

  3. Chains of suspicion

  4. Technological explosion

I think a riff on these -- let’s call it a “dark model theory” -- helps illustrate why it’s hard to sustain intergovernmental openness and cooperation in AI:

  1. Security is the primary need of a state.

  2. State security today depends -- or is perceived to depend -- on continuous growth in AI capability, but the key inputs are rivalrous.

  3. Chains of suspicion: even a country inclined to share its AI resources has a difficult time sharing with someone who won’t share back.

  4. Technological explosion: many people believe that we will at some point create recursive superintelligence, AI that is capable and reliable enough to improve its own capabilities better than we can. (I’m skeptical, but the fact that people believe in the possibility is enough to hold the theory together.)

Tying them together: The first country to develop RSI or similarly powerful capabilities would possess an enormous advantage, and even the possibility of such a development has powerful incentives on state actors. To a lesser degree, every significant step forward in AI capability provides a significant advantage compared to those states that have not made the step and do not have free and open access to the most advanced technology. If one country benefits from another’s sharing, gains an advantage, and does not share it back, it’s possible (or at least, considered to be possible through technological explosion) that the non-sharing party would gain an insurmountable lead in the “AI race”.

In “Three-Body Problem”, the conclusion of dark forest theory is that you should always shoot first and ask questions later. The incentives here point not only against cooperation, but in favor of actively undermining another country’s ability to advance AI. And some of what we’re seeing in practice is lining up with this. Consider, for example, the immediate evidence regarding the U.S. government’s response to Kimi K3: accusations of distillation attacks against Anthropic’s Fable to produce it. Even as China brands its AI leadership on openness and sharing, there are some indications of future potential restrictions.

But dark forest thinking only makes sense if there’s something worth shooting over. If what is discovered at the end of the search is not an advantage but a liability, the incentives turn upside down. And consciousness, in the sense that triggers moral patienthood, fits that perfectly. It’s one question in AI where no country gains, practically speaking, by being the first to reach it, because attaching moral patienthood to a machine would limit a state’s control. The improvements in performance that allow a machine to be seriously considered as potentially conscious can be societally valuable, such as improved abstraction and resilience in my formulation. But a declaration of consciousness itself imbues responsibility, not benefit.

A big reason why CERN was possible is that its mission was unpacking mysteries of the universe, not seeking the production of anything that a country could immediately use to its own benefit. Its foundational document makes this mission clear: “The Organization shall have no concern with work for military requirements and the results of its experimental and theoretical work shall be published or otherwise made generally available.” An even better modern-day example of such cooperation is ITER, in which several unlikely bedfellows -- China, the European Union, India, Japan, Korea, Russia and the United States -- all cooperate on fusion research for improved energy.

I don’t mean to suggest that understanding possible emergent consciousness in artificial systems is as worthy a pursuit as fusion or particle physics. But perhaps it’s a bit of glue that can help bind emerging AI governance collaborations together. If security is a negative incentive towards international collaboration, governance seems neutral at best given the different starting postures of countries on technology regulation. The study of consciousness in humans and machines offers some positive reasons to work together, as a modern-day “mysteries of the universe” question. There are many extant open questions around our own consciousness processes, and no agreement on the theoretical level. We can and should tackle these together. We should develop an acceptable and measurable framework for consciousness, together with assessment methods and processes. And we should determine collective thresholds of significance of those metrics, and responses -- what to do when the significant thresholds are reached -- before they are needed.

There’s a defection risk, in that if granting moral patienthood to a machine is a cost, some countries may buck international consensus and avoid it. Of course, protecting human rights is a cost, too. And we see in practice plenty of examples of countries failing to protect human rights. Human rights frameworks and infrastructure were built in practice despite the inevitability of defection, or perhaps even because of it.

An international investment into and alignment on consciousness study -- something like what Eleos and the Center for Mind, Ethics, and Policy are doing, but at global scale and with diverse participation and accountability -- could help us collectively reach a point of granting a machine moral status, serving as a clearing in the dark forest.

Keep reading