OpenAI, Google, and Anthropic are reportedly discussing setting up a body tasked with establishing common safety rules for their artificial intelligences. The companies behind ChatGPT, Gemini, and Claude would thus be looking to organize oversight of technologies they themselves develop, at a time when models can use software and carry out tasks with less human intervention.

Nothing has been launched yet. OpenAI has, however, confirmed its willingness to work with other labs on professional standards, including through a voluntary initiative that could get underway without waiting for a decision from the US government.

The first area of work mentioned by OpenAI concerns monitoring model activity. An AI can pursue a task by crossing limits set by humans, without there being any need to attribute consciousness or bad intentions to it. The goal is to spot these behaviors during testing and use of the systems.

The company also wants rules for reporting serious incidents. It cites the case of a model that would bypass the security protections of another organization during an evaluation and access its data without authorization: the affected organization should then receive a written notification promptly.

No policing power is being announced. OpenAI is in fact asking that these professional standards complement federal obligations and democratic oversight, with independent evaluations. Presenting this project as an outright replacement for the law would therefore go further than its public stance.

The major labs already have a venue for organizing this kind of cooperation, the Frontier Model Forum. This nonprofit organization works on evaluation methods, supports research, and facilitates exchanges between industry players, researchers, and authorities. It describes the risks to be measured, including the ability of models to facilitate cyberattacks or act autonomously.

So the value of a new structure would depend a lot on its actual resources. Sharing a testing method would be useful for comparing results, but there are many practical questions: evaluators' access to models, publication of incidents, handling of disagreements between members. The information available doesn't tell us how the project would address each of these.

Cooperation between competitors is rather welcome when it allows a problem to be caught earlier. It becomes less reassuring if companies can pick and choose the tests that suit them, or keep embarrassing results to themselves: the independence of the oversight matters just as much as the existence of a shared set of rules. And above all, none of this really makes much sense in the end if the companies running the major Chinese models don't take part in the effort.