Gemini 4 Argon Can Write a Million Tokens, and Almost Nobody Can Use It Yet

Google just announced Gemini 4 Argon, a model that can write up to a million tokens in one response. That’s a genuinely big number, and I’ve been looking forward to this kind of headroom for agents for a long time. There’s a catch, though: almost nobody can use it yet. Let’s look at what Google says Argon can do, how the rollout works, how to read the benchmark claims, and what you can do today to be ready when it opens up.

A long ribbon of glowing code and text unspooling from a single prompt window across a dark workspace, with a small locked gate in the foreground

What Google announced

On September 30, Koray Kavukcuoglu, SVP at Google DeepMind and Google’s Chief AI Architect, introduced Gemini 4 Argon on the Google blog as “our new frontier model.” Notably, the name carries no Pro or Flash label.

Google pitches Argon for long, multi-step work: real-world software engineering, enterprise knowledge work like legal and finance, cybersecurity defense, and creative writing. According to the post, “thousands of Googlers” already use it internally.

The headline capability is output length. Google says it’s expanding the output token limit “to an industry-leading 1M tokens, up from the previous 64K tokens.” Google hasn’t disclosed the model’s size, parameter count, or input context window.

The examples Google shared

These are all vendor claims. Google says Argon agents are migrating C and C++ codebases to Rust, scaling from tens of thousands of lines in libraries like re2 and libgav1 up to more than 800K lines for the Fuchsia Zircon kernel. Google notes those big rewrites are still going through auditing, emulation testing and review before production.

For libgav1, Google’s open source video decoder, the post says Argon agents replaced 32K lines of SIMD code in an existing Rust port, producing a memory-safe decoder that “runs 2.7x faster than the Rust port.” Google also says a team of Argon agents freed more than 300 TiB of memory across its data centers, and that on one quantum computing problem it “beat the published baseline by 40% in a matter of minutes.”

Who gets it first

Access is staged. Argon is rolling out first to trusted cyber defenders through Google’s Fairwind Program, a limited access program for governments and trusted partners that Google launched on September 2, along with a group of trusted testers. Google also says it’s “actively engaged in the U.S. government’s voluntary process for pre-release model access.”

Next in line are paid Gemini API customers and Google AI Ultra subscribers. Google says it wants to reach developers, enterprises and consumers “as soon as possible,” and it hasn’t given a date. When I checked the Gemini API changelog, the latest entry was September 22, and Argon wasn’t listed.

Pricing is already published. Argon will launch at an introductory $2 per million input tokens and $10 per million output tokens, with cached input 95% off. Google didn’t say how long the introductory period lasts, and afterwards the price moves to $4 and $20. One full million-token response costs about $10 at the introductory rate, and about $20 after it.

Reading the benchmark table

Google reports some strong numbers. Argon scores 77.9% on DeepSWE v1.1, which Google calls a new state of the art. It’s the leading model on the Vals Index, ranks first on Zapier’s AutomationBench at 51.3%, scores 91.7% on LVBench for long video understanding, and ties for first on CWE-bench v1 at 68%. Google also claims leading results on Vals Finance Agent v2, Harvey’s Legal Agent Benchmark and Gray Swan’s indirect prompt injection benchmark.

When I read a table like this, I ask a few gentle questions. Who ran the evaluation? Here, Google did, and I haven’t seen a model card or independent evaluation yet. What does “tied for first” or “leading” mean in numbers, and who else is on the board? Which benchmarks match the work I actually do? A 77.9% on long-horizon software tasks is exciting, and I’d still want to see how it behaves on my own codebase.

A fair word of caution

There are reasonable reasons to hold the excitement lightly. Every number above is self-reported, and developers can’t build on Argon yet, so nobody outside Google’s early cohorts can check the claims. Some observers also see the “too capable to release widely” framing that often surrounds staged frontier launches as partly marketing. I think careful staging for strong cyber capabilities makes sense, and I also think it’s healthy to keep that view in mind until independent results arrive.

What you can do now

The good news is that preparing for long-running agents doesn’t require Argon at all.

  • Build test suites for long tasks. Pick a code migration, a multi-file refactor, or an agent loop you care about, and write the checks that tell you whether a run succeeded. When a model with this much output headroom arrives, you’ll be able to measure it on your work in an afternoon.
  • Plan for longer outputs and their cost. Think about streaming, checkpointing, and how you’ll review hundreds of thousands of tokens of changes. Put the $10 to $20 per full response into your budget models now.
  • Watch for the model name. Keep an eye on the Gemini API changelog, AI Studio and Vertex AI.
  • Keep building with today’s Gemini models. Everything you learn about prompting, tools and evaluation will carry over.

If you teach, this launch makes a lovely case study. Students can compare a staged frontier release with a general launch, talk through who gets access first and why, and practice reading a vendor benchmark table with a critical, curious eye.

I’m genuinely excited to try Argon when it opens up. Until then, the best preparation is the unglamorous kind: good tests, clear budgets, and a habit of measuring for yourself.