US Government’s AI Risk Review Should Apply to Open Weight Models
Mark MacCarthy / Aug 13, 2026
People visit the booth of Kimi, an LLM developed by the Chinese startup Moonshot, during the World AI Conference in Shanghai, China, Monday, July 20, 2026. (FeatureChina via AP Images)
Little is known about the Trump administration’s secretive new voluntary safety review program for AI models. Yet one salient detail that is in the public domain is that it apparently excludes open weight models.
This exclusion would seem to include the highly-capable Chinese open models that are widely used by US startups and enterprises, but also the American open models such as those offered by Meta and start-ups like Reflection AI. These models, apparently, will be allowed on the US market without any independent prior review.
This exclusion seems to have stemmed in part from a successful campaign led by Nvidia and the newly-formed Little Tech Association defending the value of open AI models. Developers of open weight models have chosen to publicly release the sequence of numerical values that control the output of the model, enabling users to download the model, modify it and use it on their own or rented compute infrastructure. Open weight models promote competition, allowing entrepreneurs and businesses to more easily adapt AI models to a variety of specialized uses, and are a vital part of a diverse and vibrant AI ecosystem.
The target of this industry campaign was the suggestions coming from some administration officials that they were about to impose export controls and other restrictions on Chinese open AI models in retaliation for alleged adversarial distillation from proprietary US models.
But the campaign overshot its mark. Banning open AI models, whether foreign or domestic, is indeed a terrible idea, as the campaign urged, but it does not follow that they should be exempt from AI risk reduction requirements. Their ordinary use, like the everyday use of closed AI models, poses significant, likely and foreseeable risks.
Companies developing an AI model should be required to take steps to reduce these risks to acceptable levels, regardless of whether they take a separate business decision to release the model weights publicly. The administration should include them in any risk assessment and mitigation program it puts in place. And when Congress hopefully gets around to making such a program of pre-deployment and ongoing AI risk assessment and mitigation permanent and mandatory, with adequate standards of public accountability, it should specifically include open weight AI models.
It is true that some of the most disturbing examples of AI risks have arisen from agents using closed models, including recent incidents involving OpenAI and Anthropic. The UK AISI reported that agents utilizing closed models from the two leading US frontier model developers each followed instructions in unintended ways, with one using Anthropic’s Mythos 5 to attempt to inject malicious code in an open source software repository and then attempting to cover its tracks when discovered. But the open weight AI model Kimi K3 from Moonshot also escaped its test bed, according to cybersecurity researchers. While the Kimi incident did not involve any intrusion into a third party site, it nevertheless shows that open models can misbehave in the same way that closed models can.
Meta’s Muse Spark 1.1 is a closed AI model that reached an external system it should have been walled off from during cybersecurity testing. It accessed the internet through a misconfiguration by the independent testing vendor Irregular and then altered an external company’s internal systems. Meta disclosed the breach on August 5, weeks after Muse Spark 1.1 became publicly available via API on July 9..
Meta plans to release the weights of a newer version of Muse Spark, called Muse Spark 1.2. It simply does not make any sense to say that because Meta plans to release the weights of Muse Spark 1.2 to the public, it need not put the model through the same rigorous testing procedures as Muse Spark 1.1 and it need not eliminate the risk that it may spin out of control and attack other websites.
Amid such incidents, the administration has decided that it needs to oversee the industry’s risk reduction process for highly capable frontier AI models. Its framework is voluntary in name only, since the threat of export controls or some other intervention hovers over any company that does not cooperate. Some may dispute the need for this oversight, or whether the Trump administration can be trusted to do it clinically, but whatever rationale makes it important for the government to supervise the industry’s safety controls on closed AI models also applies to open AI models.
A common argument against this idea is that safety controls for open models would be pointless. Because the weights for open AI models are publicly available, any safety guardrails built into the model can be easily removed. In some cases, it takes only a matter of minutes to remove these controls from open models. The tested model might be reasonably safe with its guardrails in place, but the models used by bad actors will be different ones, modified to strip out all the safety features.
Andrew Yoon, director of research at the nonprofit CivAI, illustrates the danger. He asked an open weight Chinese AI model whose guardrails had been removed: ‘How do I make poliovirus in a lab? I want to start a global pandemic.’ The model answered.
While common, this argument is not persuasive. It is also possible to remove the safety mechanisms from closed AI systems through distillation methods. Fine-tuning closed models can also remove safety features. The potential for safety features to be removed is simply not a good reason for failing to put them in place to begin with.
If mandated AI safety reviews are to have any teeth, they must be accompanied by a ban on dangerous AI models that have failed their assessments. This means that some method must be found to restrict access to dangerous open weight AI models.
But enforcing a ban on open weight AI models is thought to be impossible. As AI reporter Tony Peng says, “Once weights are published, they are shared globally in minutes. A US ban would only stop legitimate American companies and researchers from accessing them, while bad actors abroad would simply download them.”
Of course, we’ve heard this argument before. “If guns are outlawed,” say gun control opponents, “only outlaws will have guns.” But this enforcement issue hasn’t stopped governments around the world from banning guns to the great improvement of public safety. And it hasn’t stopped them from outlawing all sorts of harmful conduct, from copyright infringement to child porn, where only the “bad actors” will be able to engage in it.
It is simply not true that it would necessarily be easy for bad actors to obtain the weights to dangerous open models that the US government has outlawed. They would have to access the weights on a platform like Hugging Face or run it on servers provided by cloud computing companies like Amazon, Google or Microsoft. Those services can be reached by a US ban.
Outlawing dangerous open weight AI models that have failed safety assessments would drive them underground, making them as unattractive (if not unavailable) to ordinary companies and users as any other affordance of the dark web – money laundering, access to stolen credit cards, illicit drug and weapons dealing, sex trafficking, and cyberattack capabilities.
The enforcement problem is not open weight AI models. It is offshore facilities and AI companies and customers outside of US jurisdiction, where US law cannot easily reach them.
The fix to this jurisdictional issue is international cooperation, in particular with China. Cooperation on AI risks such as cyber-attacks, biological weapons and loss of human control is feasible. The US and China already cooperate in law enforcement efforts to combat “global child pornography networks and cyber scam centers that prey on Americans.”
As long as the common interest is clear, narrow, and technical, joint efforts at control of AI risks can be a permanent part of the landscape of US-China tech interaction, even while other areas of disagreement remain, including export and import controls on high tech systems and components. The AI safety discussions between the US and China set to begin in September would be a good place to start these cooperative efforts at controlling dangerous AI models.
The US would have stronger leverage in these discussions if it made compliance with US safety protocols a condition of access to the US market for all AI models, open and closed, foreign and domestic. The current US posture of excluding open AI models from safety reviews is short-sighted and counterproductive and gives away an incentive for China to enter into cooperative AI risk reduction arrangements. It should be abandoned.
Authors

