Why the US Blocked Anthropic’s Claude: AI Security Risks Explained

AI security risks - Why the US Blocked Anthropic’s Claude: AI Security Risks Explained

Introduction: AI Security Risks in Focus

The recent US government action against Anthropic’s latest Claude AI models, Fable 5 and Mythos 5, has sparked intense debate about AI security risks and the ongoing struggle to regulate powerful artificial intelligence systems. On June 12, Anthropic suspended access to these models following an “export control directive” from the US government, restricting their use to US nationals. This move underscores growing concerns around the potential misuse of advanced AI, especially as these models become increasingly capable—and increasingly difficult to control.

Why Were Anthropic’s Claude Models Shut Down?

Mythos 5, Anthropic’s newest and most powerful “frontier” model, was initially released only to select organizations—primarily US tech firms—to help patch vulnerabilities in critical digital systems. Fable 5, a version with added safeguards to prevent cybersecurity misuse, was made available to the public. Yet, within days, both models were removed from access due to fears over AI security risks.

The US government has not publicly explained the details behind its directive, but Anthropic suspects the move was triggered by a jailbreak—an exploit that allowed users to bypass Fable’s built-in safety mechanisms. These safeguards are designed to classify requests as safe or unsafe, redirecting risky queries to a less capable model. However, the government reportedly grew concerned that these defenses could be circumvented, potentially enabling malicious actors to leverage the AI for cyberattacks.

Jailbreaks and the Challenge of AI Guardrails

Guardrails for large language models like Claude are far from foolproof. They largely depend on the model’s ability to interpret a user’s intentions—a task complicated by the creativity and persistence of online communities dedicated to finding workarounds. The so-called “Undersphere” of hackers and researchers continuously tests and breaks AI safety barriers, making AI security risks a moving target for both developers and regulators.

Within just 48 hours of Fable’s release, a researcher using the pseudonym “Pliny the Liberator” published what they claimed was Fable 5’s full system prompt on social media and GitHub. While the implications of this leak are not entirely clear, it drew immediate attention from those interested in probing the limits of AI safeguards.

Opaque Technology and Regulatory Dilemmas

One of the deepest challenges in addressing AI security risks is the opacity of large language models. According to experts like Oxford University’s Maximilian Kasy, these models operate as something of a “black box.” Despite being trained on vast datasets and containing billions of parameters, it’s not fully understood how they generalize so effectively—or how they can be manipulated. This lack of transparency makes it difficult for even the most experienced developers to predict, much less control, their models’ behavior.

Regulators face an even tougher challenge. Without independent access to the underlying data, infrastructure, or AI expertise, government agencies are often unable to thoroughly assess the risks posed by proprietary frontier models. The recent executive order on AI security issued by the US administration reflects this reality, signaling a shift from hands-off oversight to more intrusive demands that developers submit their models for pre-release review.

Anthropic, the US Government, and Ongoing Tensions

The showdown between Anthropic and the current US administration illustrates the complexity of AI governance. Since early 2025, tensions have flared over issues ranging from AI regulation and semiconductor policy to Anthropic’s refusal to allow its models to be used for domestic surveillance or autonomous weapons. The Pentagon’s threat to label Anthropic a “supply chain risk” would have forced government contractors to cut ties, intensifying the conflict.

Anthropic claims that the research behind the US export control directive originated with Amazon engineers—highlighting the tangled web of competition and collaboration among big tech firms in shaping AI policy.

The Global Stakes of AI Security

The uncertainty surrounding AI security risks extends far beyond US borders. A global survey last year found that people in 25 countries are more than twice as concerned about AI as they are excited by it. The rapid pace of technological change means that regulations often lag behind, while technical safeguards can be bypassed by determined actors.

Experts argue that a truly effective governance framework must be global, inclusive, and built on mutual trust. So far, the US administration has struggled to deliver such a system—leaving both industry and the public in a state of uncertainty about the future of AI safety.

Conclusion: Navigating the Future of AI Security

The saga of Anthropic’s Claude models and the US export control directive highlights just how difficult it is to manage AI security risks in an era of rapid technological advancement. As AI grows ever more powerful, neither regulation nor technical guardrails alone can guarantee safety. What’s needed is a proactive, collaborative approach that anticipates failures and adapts quickly—a challenge that will define the next chapter in AI development and governance.


This article is inspired by content from Original Source. It has been rephrased for originality. Images are credited to the original source.

Subscribe to our Newsletter