OpenAi founder Sam Altman speaks during the G20 Innovation Ministerial on September 2, 2026 in Chapel Hill, North Carolina.

Sean Rayford | Getty Images News | Getty Images

OpenAI on Monday posted a set of proposals for safety and security in the development of frontier artificial intelligence with a heavy focus on alignment research and a computing technique known as recursive self-improvement, or RSI.

“Navigating this transition safely requires alignment research to keep pace with these capabilities so that the systems we and others build remain aligned with human values and under human control,” the company said in a blog post.

OpenAI called for international cooperation to develop frontier standards and recommended building on the work of existing AI safety institutes around the world.

The ChatGPT maker said these technical standards should focus on frontier AI models and developers, as well as benefit-risk management for automated AI researchers, which includes RSI.

RSI has excited AI developers over its potential to create foundation models that can upgrade themselves without human involvement.

But advancements within RSI have led some technologists to raise concerns that foundation model makers could lose control of the underlying technology or fail to account for potential unintended consequences as the AI systems become more complicated and ubiquitous across the Internet.

“Fully autonomous RSI is not happening today, and we should not pursue it unless and until it can be done safely,” the OpenAI blog post said. “Done without appropriate care and caution, RSI could result in humans losing practical control over AI development, unable to provide oversight on research processes they no longer understand.”

The OpenAI blog post mentioned the Hugging Face agent hack, which did not involve the RSI technique, as a kind of “preview of the kinds of risks that could become much more severe without robust safeguards and alignment.”

Last week, rival Anthropic rolled out its own ideas for the safe development of frontier AI models, a response to the recent chorus of warnings about AI’s threat to humanity from industry researchers. Jacob Coxon, who has worked at both Anthropic and OpenAI, ignited a global debate when he announced his resignation nearly two weeks ago and said the companies were “gambling with our lives.”

In the aftermath of recent AI-related security incidents and Coxon’s public proclamations, Anthropic CEO Dario Amodei published an essay that called for AI companies to slow the pace of their foundation model development, among other proposals.

Amodei also raised the notion of embedding third-party evaluators into their companies as a way to audit and mitigate any potential risks that their technologies could pose to society, such as turbocharging cybersecurity-related hacks or creating bioweapons.

Rival leaders like OpenAI CEO Sam Altman and Tesla and SpaceX CEO Elon Musk also publicly supported Amodei’s proposition.

But because the field of AI evaluation is so nascent, there has yet to be a uniform consensus on the basic standards and principles that would allow independent third parties to more thoroughly inspect the cutting-edge technologies beyond what they currently do.

That’s partly why a coalition of AI evaluators are urging foundation model makers to consider a set of “minimum conditions” intended to let them more deeply perform their technology-related audits and checks, including deeper access and the prevention of retribution for publishing unflattering reports.

OpenAI reports 6 new instances of 'concerning model behavior'



Source link

LEAVE A REPLY

Please enter your comment!
Please enter your name here