// Journal · Oct 06, 2026 · 6 min read

What Gemini 4 Argon means for software teams

What Gemini 4 Argon means for software teams, concretely: a model most of you cannot call yet, an output ceiling raised from 64K to 1M tokens, and an introductory price of $2 per million input tokens and $10 per million output — roughly a fifth of what the other frontier model aimed at long-horizon work costs today. Google announced it on 30 September 2026. As of 6 October there is still no gemini-4-argon entry in the public Gemini API model list, so nothing in your codebase changes this week. What should change is how you budget and structure agent runs for the next one.

What Google actually announced

From Google’s announcement, written by Koray Kavukcuoglu, SVP of Google DeepMind:

  • Argon is positioned for long-horizon workflows: real-world software engineering, enterprise knowledge work (legal and finance), cybersecurity defence, and creative writing.
  • Rollout is phased. Trusted cyber defenders get it first through the Fairwind Program, with Google engaged in the U.S. government’s voluntary pre-release model access process. Broader access starts with paid API customers and Google AI Ultra subscribers.
  • Pricing: $2/M input, $10/M output introductory; $4/M input, $20/M output after that; cached input tokens at 95% off.

The internal usage numbers are the most interesting part of the post, because they are about long agent runs rather than chat. Argon agents analysed fleet-wide profiling telemetry and applied memory optimisations across Google’s data centres, freeing over 300 TiB (an estimated 500 TiB–1 PiB in total). Argon agents are migrating C/C++ to Rust at Google, from core libraries like re2 and libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. On libgav1 specifically, agents replaced 32K lines of SIMD code with safe Rust that the compiler auto-vectorises — 2.7× faster than the existing Rust port, with identical video output.

That is the shape of work the model is sold for: not answering a question, but grinding through a codebase for hours.

Why the 1M output token limit matters more than the context window

Everyone fixates on input context. The constraint that actually hurts when you build agents is output, and Argon raises it from 64K to 1M tokens. Google’s framing: “When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.”

If you have built an agent harness, you know what 64K cost you. A long refactor or migration had to be chopped into chunks that fit the output budget, each with its own resume logic, state summary, and stitch-back step. Half the engineering in a serious agent harness is not the prompting — it is the bookkeeping that exists purely because the model had to stop and be re-invoked. A 1M output budget deletes a chunk of that bookkeeping.

It does not delete the reason for it. A ceiling is not a target: a single uninterrupted 1M-token trajectory is also a run with no checkpoints, no intermediate review, and one very expensive failure mode if it goes off the rails at token 400,000. Longer windows also do not fix the degradation we wrote about in context engineering for AI agents — more room to generate is not more room to stay coherent. Our default stays what it was: long runs, frequent checkpoints, and a human-in-the-loop review queue at every irreversible step.

What the benchmarks say about agentic and automation work

The published scores, with what each one is actually measuring:

BenchmarkScoreWhat it measures
DeepSWE v1.177.9% (state of the art)Real-world long-horizon software engineering tasks
AutomationBench (Zapier)51.3%, ranked #1End-to-end execution across core business functions
LVBench91.7% (state of the art)Long video understanding
CWE-bench v168%, tied firstRemediating security vulnerabilities
Vals IndexLeadingFinance, coding, legal and tax work, weighted by U.S. GDP contribution
Gray Swan IPILeadingRobustness against indirect prompt injection

Read AutomationBench carefully, because it is the one closest to what automation projects actually do. The best model in the world completes 51.3% of end-to-end business tasks. First place is a coin flip. If you are scoping an automation on the assumption that a frontier model will reliably drive a multi-step business process unattended, that number is your reality check — and the reason confidence thresholds and escalation paths are not optional.

The Gray Swan IPI result is the underrated one. Any agent that reads email, tickets, PDFs or web pages is ingesting text an attacker can write. Prompt-injection robustness is a load-bearing property for every automation of that kind, not a safety footnote.

What $2/$10 per million tokens costs on a real automation

Take a realistic agentic run: 200K input tokens of code and context, 300K output tokens of reasoning and diffs.

  • Argon, introductory: $0.40 input + $3.00 output = $3.40 per run
  • Argon, standard: $0.80 + $6.00 = $6.80 per run
  • With 95% cached input (same repo context, many runs): input drops to $0.02–$0.04, so the output dominates almost entirely

For comparison, GPT-6 Astra lists at $10/M input and $50/M output — the same run costs about $17. On list prices, Argon’s introductory rate is roughly a fifth of that, and its standard rate around 40%.

Two things follow. First, output tokens are 5× the price of input, so the 1M output ceiling has a price tag: a maximal single generation is $10 introductory, $20 standard, in output alone. “Let it think as long as it wants” is now a budgeting decision, not a technical one. Second, with caching at 95% off, the economics favour many runs against a stable context — exactly the pattern for a repo, a document corpus, or a policy manual that barely changes between runs. Design your prompt cache boundaries before you design your agent loop.

What Gemini 4 Argon means for software teams right now

Argon is not generally available, so the useful work this week is preparation, not migration:

  1. Keep the harness model-agnostic. Anything that hard-codes a provider’s quirks will be rewritten twice before year end. Swapping models should be a config change.
  2. Instrument token spend per run now. You cannot evaluate a $3.40-per-run model against your current one without knowing what your current one costs per completed task — not per call.
  3. Measure completed tasks, not benchmark scores. 51.3% on AutomationBench says nothing about your process. Build a small eval of real cases from your own backlog and run every candidate model against it.
  4. Do not loosen your guardrails in anticipation. Nothing in this release removes the need for least-privilege credentials and human approval before irreversible actions.

Havoric is an AI automation and web development agency — we automate repetitive manual processes and build the web and mobile apps around them. The question teams bring to us as an AI agency is usually “which model should we use?”, and the honest answer rarely changes with a launch: the model is the cheapest component to swap and the smallest source of risk in the system. Our AI automation scoping will treat Argon as a strong candidate for long-horizon engineering work the moment it has a model ID we can call and a quota we can rely on. Until then, the thing worth building is the harness around it — the checkpoints, the evals, the review queue — because that part outlives every model you point it at.