Our Blog

Blog Index 

Google Launches Gemini 4 Argon: Frontier AI Built for Cyber Defense, Rolling Out First to Trusted Defenders

Posted on 1st Oct 2026 12:05:53 in Artificial Intelligence, Machine Learning

Tagged as: Google, Gemini 4 Argon, Google DeepMind, AI models, cybersecurity, frontier AI

Google introduced Gemini 4 Argon on Wednesday, September 30, calling it the next era of its frontier intelligence. The model is the company's most capable AI system to date, and Google says it delivers frontier performance across three demanding areas: real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.

But most people cannot use it yet, and that is entirely deliberate. Argon is rolling out first to a set of trusted cyber defenders through Google's Fairwind Program, a channel built for governments and vetted partners that need the company's strongest cybersecurity capabilities. Google says it is actively engaged in the U.S. government's voluntary process for pre-release model access while it gradually expands access. The broader AI market will have to wait its turn.

The launch lands in a heavy week for the industry. It comes a day after OpenAI's DevDay conference and just after Sundar Pichai, Google's chief executive, signed a voluntary AI safety accord with President Donald Trump and other major technology leaders at the White House.

What Gemini 4 Argon Brings to the Table

According to Koray Kavukcuoglu, Google's chief AI architect and the newly appointed head of Google DeepMind, Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work, and cybersecurity defense. The model is built to sustain deep reasoning over long-horizon tasks, and it attacks the problem from several directions at once:

  • A one-million-token output limit — up from 64,000 tokens in the previous generation. Google calls the limit industry-leading and says it gives the model room to think deeply and generate hundreds of thousands of tokens in a single trajectory, which matters for problems that cannot be solved in one short answer.
  • State of the art on DeepSWE v1.1 (77.9%), a benchmark that measures performance in real-world, long-horizon software engineering tasks.
  • First place on the Vals Index, which measures economic impact across finance, coding, legal, and tax work, with sectors weighted by their contribution to the U.S. economy. Google also reports leading results on Vals Finance Agent v2 for multi-step financial research and Harvey's Legal Agent Benchmark for legal research and drafting.
  • Number one on Zapier's AutomationBench (51.3%), which measures end-to-end execution across core business functions.
  • State of the art on LVBench (91.7%) for long video understanding, part of the model's push into knowledge work that depends on visual comprehension.

Google says Argon can analyze charts professionally, identify fine details in long-form video, and work through a series of documents in one session. Tulsee Doshi, Google's Gemini model product lead, described Argon to CNBC as an incredibly well-rounded model that excels at running long, multi-step tasks. On Artificial Analysis's Intelligence Index, an independent benchmarking service, Argon matches OpenAI's GPT-6 Astra at roughly 60 percent of the cost per task at current discounted prices.

Built for Cyber Defense

The sharpest edge on Argon is defensive security. Google trained the model to be highly capable at cybersecurity defense, and says it can autonomously find, validate, and patch critical software vulnerabilities. On CWE-bench v1, which evaluates a model's ability to remediate security flaws, Argon tied for first place with a score of 68 percent, building on the performance of Google's 3.8 Flash Cyber model released earlier this month. It also leads Gray Swan's Indirect Prompt Injection benchmark, making it Google's most resilient model yet against attacks that try to hijack a model's behavior through malicious instructions hidden in content.

In an early demonstration, Google says Argon uncovered a critical vulnerability in healthcare software used by hospitals around the world, exposing sensitive personal information that previous frontier models had missed. Security firm Wiz is already using the model through its Scan for Good initiative, a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures.

For trusted defenders and Google's own internal security teams, the company will release Argon without cyber guardrails, so they can use its full frontier-level cybersecurity capabilities. That is a notable amount of trust placed in a model, and it cuts both ways: the same capability that finds and fixes flaws could be pointed at systems that were never meant to be tested. It is precisely why access is gated in the first place.

Already at Work Inside Google

Argon is not waiting for the public rollout to prove itself. Google says the model is already powering internal workflows, with thousands of employees using it for specialized coding tasks, deeper research, and writing quality. Three uses stand out:

  • Quantum algorithm optimization: Argon is helping Google's quantum computing researchers optimize the spacetime resources, measured as qubits times gates, of subroutines that bottleneck important applications.
  • Data center memory efficiency: a team of Argon agents analyzed fleet-wide profiling telemetry and autonomously identified and applied memory optimizations across Google's data centers, freeing more than 300 tebibytes of memory once rolled out, with estimated total savings between 500 TiB and one pebibyte.
  • Large-scale codebase migrations: Argon agents are helping migrate C and C++ codebases to the memory-safe Rust language, scaling from tens of thousands of lines in core libraries such as re2 and libgav1 up to more than 800,000 lines for the Fuchsia Zircon kernel. For libgav1, Google's open-source video decoder, Argon replaced 32,000 lines of SIMD code through profile-guided experiments and produced a Rust port that runs 2.7 times faster than the earlier port with identical video output.

Why the Gated Rollout

Google is explicit that this is a safety-first release schedule. Kavukcuoglu says the company is limiting access at first so it can make sure the model is not misaligned, and that it will strengthen critical frontier safeguards before a broader rollout. Doshi framed the phased approach as a confidence builder: starting the rollout this way gives Google more confidence, she told CNBC, while also putting a model trained for cyber defense into the hands of defenders as soon as possible.

Before broad availability, Google says it is hardening four areas: defending against misuse for cyber or chemical, biological, radiological, and nuclear attacks, with red-team testing by internal and external teams; defending against indirect prompt injection attacks; monitoring for misalignment by watching the model's chain of thought and actions and stopping execution when necessary; and hardening the sandboxed environments used to test and evaluate the model.

The gated pattern is becoming industry standard for the most capable systems. OpenAI and Anthropic have also released their strongest models to restricted sets of partners first, citing cybersecurity concerns. Argon also comes just weeks after Google's own 3.8 Flash Cyber release, and a day after OpenAI launched a new agent and model of its own while holding back a planned flagship over safety worries. Frontier AI is now released less like a product and more like a controlled substance.

Pricing and Availability

When Argon does reach the broader market, pricing will be aggressive. The introductory rate is $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95 percent off. After the introductory period, those rates rise to $4 per million input tokens and $20 per million output tokens. For comparison, OpenAI's GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens, and Argon's one-million-token output ceiling is several times Astra's 128,000-token output limit.

Wider access will start with paid API customers and Google AI Ultra subscribers, then extend to developers, enterprises, and consumers as soon as possible, Google says. The company will keep gathering feedback from early testers and iterating on guardrails before each expansion step.

What This Means

Gemini 4 Argon resets the frontier conversation. Google claims records in real-world software engineering, a tie for first in cybersecurity, and leadership on economic-impact benchmarks covering finance, legal, and professional work, all while undercutting rivals on price. On cybersecurity benchmarks, Argon ties OpenAI's GPT-6 Astra and xAI's Grok 4.7, and it edges ahead of OpenAI's GPT-6.1 Sol on Artificial Analysis's composite index by a single point.

For developers and businesses watching from the sidelines, the practical takeaways are straightforward: expect API access to arrive in phases, with paid customers first; the introductory pricing window is worth planning around for budget-sensitive teams; and the one-million-token output limit opens the door to agentic workflows that previously were cut short mid-task. Security teams at large organizations should also note the direction of travel, as models trained to find and patch vulnerabilities autonomously could reshape how software audits are done.

The larger story is the industry's new release ritual. Capability announcements now come bundled with safety frameworks, government processes, and limited-access programs, and Google has leaned fully into that playbook. The real test will come when independent evaluators can stress-test Argon's claims and when the velvet rope comes down for everyone else.

Sources

whatsapp me