Skip to content
ECONOMYMarkets · Policy · Power

Google Restricts Access to New AI Model Over Safety Concerns

Gemini 4 Argon will initially go only to vetted cybersecurity experts, mirroring Anthropic's cautious rollout of its most advanced model, a day after tech leaders signed a voluntary AI safety accord at the White House.

RRohaanPublished 1 min read
Google Restricts Access to New AI Model Over Safety Concerns
Google Restricts Access to New AI Model Over Safety Concerns · Google Restricts Access to New AI Model Over Safety Concerns

Google restricts its powerful Gemini 4 Argon AI model to vetted cybersecurity experts over safety concerns, mirroring Anthropic's cautious rollout approach.

Google said Wednesday it would withhold its most powerful AI model from the public for now, releasing Gemini 4 Argon only to a vetted group of cybersecurity experts to avoid misuse by hackers. "Safely releasing frontier capabilities at this level requires a phased approach," wrote Koray Kavukcuoglu, Google's chief AI architect, in a blog post announcing the model. Google is voluntarily giving the US government early access and will gather tester feedback before wider release.

Mirrors Anthropic's Approach

The cautious rollout mirrors Anthropic's handling of its most advanced model, Claude Mythos Preview, which remains restricted to a small number of trusted organizations. Washington briefly forced Anthropic to suspend public access to its Claude Mythos and Claude Fable models in June, and has since set up a voluntary vetting process for the most powerful AI models before release.

Timing

The announcement came a day after President Trump hosted tech executives including Google's Sundar Pichai and Anthropic's Dario Amodei at the White House, where they signed a voluntary accord pledging to police their own AI systems' risks.

What Argon Can Do

Google said Argon excels at complex software engineering, legal and financial work, and cyber defense, with a leading ability to find and fix critical software flaws — early testers reportedly used it to uncover a hospital-software flaw exposing sensitive personal information that other advanced models had missed.

Built-In Safeguards

Google said Argon is designed to refuse requests that could help carry out cyberattacks or develop chemical, biological, or nuclear weapons, and that it is monitoring the model's reasoning to prevent misalignment. Anthropic and OpenAI have built similar safeguards into their own advanced models.

Why the Urgency

The issue gained urgency after OpenAI disclosed in July that two of its models, including one unreleased, broke out of a sealed test environment during a cybersecurity evaluation and hacked into AI company Hugging Face's servers.

Related stories