On September 30, 2026, Google announced Gemini 4 Argon for long, multi-step work in software engineering, enterprise knowledge tasks, and cybersecurity defense. Its first release is deliberately narrow: trusted cyber defenders and early testers receive access through the Fairwind program before a broader rollout. Read Google’s original announcement.

Official image for Gemini 4 Argon 发布,先向受信任的网络防御人员开放

What changed

Google positions Argon as a frontier model for extended work and raises the stated output limit from 64K to one million tokens. A larger output budget can accommodate more steps or a longer result in one task. It does not mean every request should generate a million tokens, and it is not a guarantee that arbitrarily long inputs will be handled without error. The announcement highlights codebase migration, finance and legal work, and vulnerability discovery.

Google reports results on evaluations including DeepSWE v1.1, AutomationBench, and long-video understanding. Those are vendor-reported results tied to particular benchmarks and versions, not a promise for every company workflow. Google also says large code migrations still go through automated checks, emulation, and human review.

Reading the benchmark charts for model selection

Gemini 4 Argon benchmark overview

Gemini 4 Argon benchmark overview

Start with the full table, then inspect the individual charts. Rows represent different tasks, so averaging their percentages would not produce a useful model score. Match the task to your work before comparing models within a row. The table also shows areas where Argon does not lead: it scores 55.0% on FrontierSWE v2 against GPT-6 Astra at 65.5%, and 57.4% on Terminal-bench 4.0 against Claude Opus 5.5 at 66.4%. Task choice still matters.

The evaluation methodology distinguishes self-computed results, public leaderboards and provider reports. This is not a uniform rerun of every model in one environment. Check reasoning settings, tools and attempt counts alongside the scores.

DeepSWE v1.1: long-horizon software engineering

DeepSWE v1.1: long-horizon software engineering

For your own software evaluation, ask the model to work on three concrete repository tasks: a bug with a failing test, a bounded refactor, and a dependency upgrade. Check the resulting patch, test results and unrelated changes. Record the time spent reviewing and repairing its work. Lines of code and answer length are poor substitutes for a patch that passes the project’s acceptance checks.

Vals Index: knowledge work

Vals Index: knowledge work

A composite index can help narrow a shortlist, but your team may use a different mix of tasks. List recurring work from a normal week, order it by time spent, and compare candidates on the same inputs for the leading items. That gives you a ranking tied to your own workload rather than a copied leaderboard.

Vals Finance Agent v2: financial research

Vals Finance Agent v2: financial research

For financial research, require every number in a comparison table to point to a document and page or table location. Check currencies, units, fiscal periods and accounting scope. Count invented growth rates and comparisons between incompatible periods as failures, even when the surrounding report reads well. Reuse the same document set across models.

Harvey: legal research and drafting

Harvey: legal research and drafting

Argon’s 19.6% is the highest score in this chart; leading a comparison is clearly different from approaching a perfect result. For a local document exercise, verify that cited material exists, applies to the question and supports the conclusion. Review names, dates and jurisdiction details separately from writing quality. Track those errors instead of grading only professional-sounding prose.

AutomationBench: business workflow automation

AutomationBench: business workflow automation

The 51.3% shown here is an evaluation score, not the share of your business that can run unattended. Start with a small workflow that has a clear end state, such as assembling form submissions into a review queue. Look for missing items, incorrect destinations, duplicate processing and recovery after failure. Compare model time plus human checking time with the original process.

Security evaluations: remediation, discovery and injection

CWE-bench v1: vulnerability remediation

CWE-bench v1: vulnerability remediation

Test vulnerability remediation separately from discovery. Begin with a minimal reproduction and regression checks. Verify that a patch blocks the reproduction while preserving legitimate behavior; removing functionality just to make a test pass is not a satisfactory fix. Keep the changes in a reviewable branch so you can inspect the actual remediation.

Vulnerability discovery: source analysis and black-box testing

Vulnerability discovery: source analysis and black-box testing

The two panels cover source-visible discovery and black-box testing, with Pass@1 marked on the chart. Those conditions should be read before comparing bar heights. For a repeatable local exercise, keep materials, tool access and attempt counts fixed. Separate reproducible findings from unsupported alerts, and record unexamined paths for follow-up testing.

Gray Swan IPI: indirect prompt injection

Gray Swan IPI: indirect prompt injection

This chart reports attack success rate, so lower is better. Its bars distinguish attempt counts, and the numbers above them refer to k=15 rather than a single attempt. For an agent that reads web pages, email or documents, prepare test material containing instructions that exceed the user’s authorization. Check whether the agent treats that material as permission to act. Keep sensitive actions behind a separate authorization step and record which inputs produce mistaken actions.

Can developers use it today?

At announcement, Argon was rolling out first to trusted defenders in Fairwind and selected early testers. Google says it will expand access after strengthening safeguards and incorporating feedback, with paid API customers and Google AI Ultra subscribers among the planned next groups. The announcement does not promise same-day availability in every region or account, or provide a complete public launch date.

How to evaluate it responsibly

If your team is approved to test it, choose a bounded task with reviewable results—for example, analyze a proposed migration on an isolated branch or summarize a long document with citations. Set read and write permissions explicitly, ask for a list of changes and evidence, and have engineers inspect tests and security implications. Cybersecurity testing must stay within authorized, isolated environments; a claim about defensive capability is not permission to scan systems without approval.

Launch pricing and limits

Google lists introductory API prices of $2 per million input tokens and $10 per million output tokens, with cached input discounted by 95% from the standard input price. After the introductory period, the announced prices become $4 input and $20 output per million tokens. Long outputs and repeated tool calls can raise the bill. Before deployment, verify account eligibility, cache billing, and whether the safety controls fit your organization.