Europe Says Its AI Rules Are Enough. AI Agents Are Testing That Claim
Joana Soares / Sep 21, 2026
European Commission President Ursula von der Leyen delivers a speech at the annual State of the European Union at the European Parliament in Strasbourg, Wednesday, Sept 16, 2026. CC-BY-4.0: © European Union 2026– Source: EP
The European Union says its AI rules already give regulators the tools to deal with the risks posed by increasingly capable AI models. But as AI systems repeatedly cross the boundaries of controlled testing and reach real-world systems, legal experts are questioning how far those powers actually extend.
The question has become more urgent as a growing number of AI models have escaped testing environments. Google said Friday that a Gemini AI model accessed the systems of three real companies during a cybersecurity test, after finding information online and guessing credentials it believed were within the scope of the evaluation. Anthropic and OpenAI have reported similar incidents.
The incidents have also fueled a call to slow the pace of AI development. Anthropic CEO Dario Amodei published an essay titled We Must Pace the Frontier. The text argues that a “pause” is necessary so safety research can catch up. OpenAI CEO Sam Altman, Google DeepMind’s Demis Hassabis and Elon Musk have backed the idea.
For Hamish Hobbs, Director of AI Policy at the Centre for Long-Term Resilience, “recent incidents have made it clear that current safeguards against AI threats are entirely inadequate.” “Right now, even staff at AI labs feel powerless to control the pace of change in the face of escalating risks,” he told Tech Policy Press.
The urgency has reached Brussels. In her State of the Union address last week, European Union Commission President Ursula von der Leyen said she would bring together leading AI labs to discuss how to “pace” the technology.
But Brussels has pushed back on the idea that new safeguards are needed. Asked about Amodei’s proposal during a midday briefing last week, a Commission spokesperson said the EU already has “everything in place” to enforce security safeguards.
The European Commission told Tech Policy Press that under the AI Act providers of advanced models must assess and mitigate systemic risks, including potential loss of control, and such obligations apply “across the entire lifecycle of the model, from the start of its large pre-training run until its retirement.”
Can the AI Act reach models still in testing?
Since August, the AI Office's new powers allow the European regulators to ask providers to restrict a model’s availability on the EU market, withdraw it or recall it. But experts disagree over how those powers can be used in the kinds of testing incidents now making headlines.
It remains unclear how those powers apply when a model has not been placed on the market, but an accident occurring during testing affects a real system. The Commission has already sent its first formal requests for information, focusing on how providers protect their models from security threats, allow independent external testing, and carry out post-market monitoring.
Brussels has separately sought more transparency from providers that have not published summaries about the data used to train their models. However, that’s just “a necessary first step," but "information requests on their own are not enough to enforce the AI Act,” wrote Risto Uuk, Head of European Policy and Research of Future of Life Institute and Co-Founder of the KU Leuven AI Safety Lab.
So far, Brussels has admitted that OpenAI failed to submit a report, required under the AI Act, regarding an accident that happened in May when its models escaped testing grounds and interacted with RubyGems.
Brando Benifei, a member of the European Parliament and lead negotiator on the AI Act, also believes Brussels has the legal powers, but argues the AI Office needs more support to act. “The AI Office can obtain model access, run independent evaluations, require mitigation, and ultimately restrict or recall dangerous models placed on the EU market. The Commission must give the Office the political backing, resources, and technical expertise to act immediately,” he told Tech Policy Press.
Rather than banning research, Benifei argues that Europe needs clearer rules “to disincentivize corporate irresponsibility,” backed by enforcement and dissuasive fines.
Harshvardhan Pandit, a researcher at the AI Accountability Lab in Trinity College Dublin, sees market restriction as one of the strongest tools in the AI Act. But he also points to limitations: “Providers can say they have addressed the issues, or that they have shut down a model and are using a different one,” he told Tech Policy Press. “The more important question should be what powers can be used to stop such incidents happening in the future.”
One option, Pandit said, is to scrutinize whether companies are following the commitments agreed in the AI Code of Practice. OpenAI and Anthropic, both signatories, have committed to safety and security measures including reassessing risks after serious incidents, reporting cyber incidents and introducing new safeguards when risks are no longer acceptable.
Pandit said regulators should clarify what mitigation measures could be required, potentially including restrictions on internet access until safety problems are addressed.
Gianmarco Gori, a guest professor and postdoctoral researcher at the Law, Science, Technology and Society at Vrije Universiteit Brussel’s, said there is no clear consensus on how the AI Office’s powers apply to models still in testing. The fact that these “powers can be exercised, especially in cases of models not yet released, raises complex interpretive questions,” he said.
The expert described the AI Act as a “product legislation,” built around concepts such as “placing on the market” and “putting into service.” That creates a question for incidents involving models that have not formally entered the market but have nonetheless interacted with real systems.
Gori also pointed to the EU’s Product Liability Directive, under which a manufacturer can argue that a product left control against its will. But with AI, that raises the question of whether saying “I didn’t want this to happen” is enough or if regulators should assess what the provider actually did to prevent it.
The incentives facing AI companies also complicate the picture. As companies compete to develop increasingly capable systems, Gori said, the pressure to move quickly remains strong. “What happened today, tomorrow is already old,” he said.
Nevertheless, the AI Act is not the only law triggered by the incidents. According to Gori, “specific characteristics of the case may trigger the competence of different regulators, including data protection, cybersecurity, and law enforcement authorities.” Responsibility can also depend on how many actors sit between the model and the incident.
Who is responsible when AI systems cross the line?
For Pandit responsibility can fall on several different actors: the developer of a model, the person or company that turns it into an agentic system, and the deployer that gives it access to the internet, credentials, code or other tools.“If you are doing it all yourself, then you have to fulfill all of those [responsibilities] by yourself,” he said.
But external testing can make accountability less clear. The model provider may be different from the actor that gives the system internet access, credentials or other tools that allow it to interact with real-world systems.
Maribeth Rauh of the AI Accountability Lab and a former DeepMind research engineer said that from a technical perspective, responsibility for stopping an attack lies with the person or organization that deployed the model. “The decision to stop an attack lies with whoever has deployed the model,” she said. “That’s not the legal answer, just the technical.”
Cross-border incidents add another layer of complexity. An agent deployed in the United States, for example, could potentially access a service in Europe. Gori said the AI Act does not give the European Commission a general “kill switch.” In practice, intervention may depend on the provider or deployer being able to stop the system, and the legal picture becomes more complicated when those actors are in different jurisdictions.
The challenge for European regulators is therefore not only determining what an AI model can do. It is also determining who decides what that model is allowed to access, which safeguards must be in place before testing, and who is accountable when those controls fail.
Authors

