Back to all posts

cat posts/gemini-4-argon-third-2-10-model-your-azure-gap.md --category "AI & Cloud Infrastructure" --views 42

Gemini 4 Argon: the third $2/$10 model and your Azure gap

Gemini 4 Argon launched on 30 September at an introductory $2/$10 per million tokens, the third frontier model at that price in three days, but unlike Claude Sonnet 5.5 and GPT-6.1 Sol it is not in Microsoft Foundry, has no API model ID, and goes to vetted cyber defenders first. The standard price is $4/$20 with no published end date for the discount, so Azure teams in Sweden should budget on the higher figure and prepare a gateway route rather than switch now.

  • --author By Falak Mahmood
  • --date October 1, 2026
  • --read 14 min read
  • --views 42 views

Google announced Gemini 4 Argon on 30 September 2026 at an introductory price of $2 per million input tokens and $10 per million output tokens. That is the third frontier-class model to land on exactly that price in three days, after Claude Sonnet 5.5 on 28 September and GPT-6.1 Sol on 29 September. The difference for an Azure team in Sweden is that the first two are in Microsoft Foundry today and Argon is not, and will not be. Google is shipping it first to vetted cyber defenders, has published no API model ID, and has not said when the introductory price ends. Here is what the announcement contains, what it costs once the discount expires, and how a Foundry shop would reach it if the eval numbers justify a third vendor.

What Google announced

Argon is Google DeepMind's new frontier model, positioned for software engineering, enterprise knowledge work in law and finance, and defensive cybersecurity. The headline specification is output length: a single generation can run to 1 million output tokens, up from 64K on the previous generation. Google's framing is that a model with room to think and generate hundreds of thousands of tokens in one trajectory solves hard problems in a single pass instead of a loop of short turns. The input context window is not stated in the announcement.

Access comes in phases. Today the model is rolling out to trusted defenders through the Fairwind Program, Google's limited-access cyber defense scheme launched on 2 September. Google says it is also taking part in the US government's voluntary pre-release model access process. Broader availability "as soon as possible" starts with paid API customers and Google AI Ultra subscribers. Koray Kavukcuoglu, SVP at Google DeepMind, is quoted by Unite.AI saying that safely releasing frontier capability at this level requires a phased approach.

Google also lists internal results it attributes to Argon: a 40% improvement over a published baseline on a quantum algorithm optimisation, more than 300 TiB of data-centre memory freed, an 800,000-line C and C++ to Rust migration in the Fuchsia Zircon kernel, and a 2.7x speed-up in the libgav1 video decoder. Those are Google's own numbers on Google's own systems. Treat them as a capability demonstration, not a forecast for your codebase.

Benchmarks: ahead in 13 of 18, behind where it matters to many of you

Google published its comparison charts as images, so the competitor figures below are as read from those charts by VentureBeat and DataCamp. The Argon figures also appear in Google's prose. VentureBeat counts Argon leading or tying on 13 of 18 disclosed benchmarks against GPT-6 Astra and Claude Opus 5.5.

Benchmark Gemini 4 Argon GPT-6 Astra Claude Opus 5.5
DeepSWE v1.1 (agentic coding)77.9%74.1%74.2%
FrontierSWE v255.0%65.5%62.3%
Terminal-bench 4.057.4%58.2%66.4%
AutomationBench (business workflows)51.3%41.4%42.5%
Vals Finance Agent v265.4%53.5%58.6%
Harvey Legal Agent19.6%5.4%3.8%
GraphWalks 256K to 1M (long context)84.2%71.8%66.8%
CWE-bench v1 (vulnerability remediation)68.0%68.0%67.0%

Two patterns stand out. Argon's largest leads are on long-horizon business and document work: finance agents, legal agents, business automation and long-context retrieval. Its largest deficits are on FrontierSWE v2 and Terminal-bench, where Opus 5.5 leads by 9 points, and on Terminal-Bench Science where Astra leads by 10.5. If your main workload is a coding agent in a terminal, the model that tops the chart is not Argon. If it is contract review or financial analysis over long documents, the gap runs the other way.

On the security side, Google's cyber page reports 85.8% on its real-world vulnerability discovery set against 71.0% for Gemini 3.8 Flash Cyber, 70.9% on the Wiz penetration test benchmark against 58.2%, and a 0.7% attack success rate on Gray Swan's indirect prompt injection benchmark, which Google describes as the lowest among models tested.

The price sheet, with the discount and without it

Google's wording is precise: Argon "will launch at an introductory price" of $2 and $10, with cached input at 95% off, and "after the introductory period expires" the price becomes $4 per million input tokens and $20 per million output. No end date is given. The standard price is therefore Claude Opus 5.5 money, and the introductory price is GPT-6.1 Sol money. Which one you budget on decides whether Argon is the cheapest frontier option or tied for the second most expensive.

Model, list price 1 October 2026 Input / MTok Cached input Output / MTok In Foundry
Gemini 4 Argon, introductory$2.00$0.10$10.00No
Gemini 4 Argon, after introductory period$4.00$0.20$20.00No
GPT-6.1 Sol, Global Standard, up to 272K$2.00$0.10$10.00Yes, EU Data Zone at +20%
Claude Sonnet 5.5, Global Standard$2.00$0.20$10.00Yes, no EU Data Zone
Claude Opus 5.5, Global Standard$4.00$0.20$20.00Yes
GPT-6 Astra, Global Standard$10.00not listed$50.00Yes, Limited Access

Google has a recent precedent for how introductory pricing ends. The Gemini API pricing page lists Gemini 3.8 Flash at $0.75 input and $3.75 output "through December 31, 2026" and $1.50 and $7.50 "starting January 1, 2027", a clean doubling on a published date. Argon's doubling is published. Its date is not. Write the $4/$20 figure into any business case that runs past this year and treat the $2/$10 period as upside.

Gemini also bills context caching differently from Anthropic and OpenAI. There is no cache-write premium; instead you pay a storage price per million tokens per hour that the cache stays alive, $0.50 for 3.8 Flash during its introductory period. Argon's storage rate is not published. For the 60-turn, 150K-prefix agent session we have used across the Sonnet 5.5 and GPT-6.1 Sol comparison and the Opus 5.5 migration math, 9M cached input tokens and 120K output tokens come to $2.10 on Argon's introductory price and $4.20 on the standard price, before the storage line. GPT-6.1 Sol runs the same session for $2.55 on Global Standard, including cache writes. The one-cell difference that decided last week's comparison, the $0.10 cache read, is a tie between Sol and introductory Argon.

The 1M output token line is a budget control problem

A model that can emit a million tokens in one call can bill $10 for that call at the introductory rate and $20 later, and Google's own argument for the feature is that the model should be allowed to run long. That is a different cost profile from a 128K cap, which is the limit on GPT-6.1 Sol and on Claude in Foundry. Three controls belong in the design before the first production request.

  • Set an explicit output cap per call type. A classification endpoint has no business with a six-figure output budget. Decide the cap per route, not per model.
  • Put a token gateway in front of it. Azure API Management's llm-token-limit policy enforces tokens per minute and a quota per hour, day, week, month or year per subscription key, and returns 429 on the rate limit and 403 on the quota. It works on any OpenAI-compatible backend, including Gemini's.
  • Log reasoning tokens separately. DataCamp notes Google has not said whether Argon's reasoning tokens bill at the output rate. On Gemini 3.8 Flash the pricing page says output price includes thinking tokens. Assume the same until the model page says otherwise.

Where Argon is not, and how an Azure team would reach it

Microsoft Learn's list of Foundry Models from partners and community, updated 21 September, names Anthropic, Cohere, Meta, Microsoft, Mistral AI and NTT Data. There is no Google section, and there has never been one. On Google's own side, the Gemini Enterprise Agent Platform release notes, current to 29 September, record Claude Sonnet 5.5 and Opus 5.5 arriving in Model Garden but no Argon entry, and the locations page lists Gemini 3.8 Flash Cyber in the Europe regions but not Argon. DataCamp confirms no API model ID had been published as of 30 September. So "reach it" today means "be ready for it", and there are two routes.

Route 1: Gemini Developer API through Azure API Management

Microsoft documents importing Google's OpenAI-compatible Gemini endpoint into API Management as a Language Model API. The base URL is https://generativelanguage.googleapis.com/v1beta/openai, the key goes in as a Bearer token stored as a named value, and the usual AI gateway policies apply: token limits, token metrics, semantic caching, content safety. Your application code keeps calling an OpenAI-shaped chat completions endpoint on your own APIM hostname, and the model name is the only thing that changes when Argon gets an ID.

// Same client, same shape, different backend behind APIM.
// Model ID is a placeholder: Google has not published one for Argon yet.
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://tsp-apim-swc.azure-api.net/gemini",   // APIM route to generativelanguage.googleapis.com/v1beta/openai
  apiKey: process.env.APIM_SUBSCRIPTION_KEY,               // APIM key; the Google key lives in APIM as a named value
});

const res = await client.chat.completions.create({
  model: "gemini-4-argon",          // replace with the ID Google publishes
  max_tokens: 8000,                 // cap per route; the model allows up to 1M
  messages,
});

console.log(res.usage);             // prompt_tokens, completion_tokens: feed llm-emit-token-metric

The catch is residency. The Gemini Developer API is a global service with no in-region processing guarantee, and Google's own docs direct customers with residency requirements to the enterprise platform instead. For a Swedish team this route is fine for evaluation and for workloads with no personal data. It is not the production path for anything a DPIA touches.

Route 2: Gemini Enterprise Agent Platform with the EU endpoint

Google's enterprise surface, formerly Vertex AI, exposes Gemini models on regional endpoints and on two multi-region endpoints that Google says keep machine-learning processing of customer data inside a jurisdiction: aiplatform.us.rep.googleapis.com and aiplatform.eu.rep.googleapis.com. The Europe regional list includes Finland (europe-north1), Belgium, the Netherlands, Frankfurt, Warsaw and others. It does not include Google's Stockholm region. The locations page also warns that ordinary endpoints do not guarantee residency and that the global endpoint should not be used if you have processing-location requirements. This is the equivalent of Foundry's EU Data Zone, and it is where Argon will have to appear before a regulated Swedish workload can use it. Private connectivity to the multi-region endpoints needs Private Service Connect, not Private Google Access.

Running this from Azure means a second cloud account, a second identity boundary, cross-cloud egress on every request, and a second set of logs to pull into Sentinel or Log Analytics. None of that is exotic. All of it is work that Sonnet 5.5 and GPT-6.1 Sol do not require, which is the real price difference between a model in Foundry and a model outside it.

The Fairwind door, and the Foundry equivalent

For most readers the Fairwind Program is a footnote. For a Swedish operator in energy, telecoms, healthcare or financial services it is the only door that is open today. Google's program page prioritises governments and national cyber authorities, critical infrastructure operators in exactly those four sectors, and maintainers of core technology platforms. Applicants are vetted for a track record of ethical operation, submit a form, and agree to conditions: phishing-resistant MFA, access limited to internal security, incident response or penetration testing teams, no sharing or resale of access, defensive and research use only. Google counts more than 650 participating partners globally and names no geography. Members receive Argon without cyber guardrails, plus the CodeMender patching agent, delivered through Gemini Enterprise with zero data retention or through Google Cloud. Organisations that do not qualify can run CodeMender with publicly available models.

Azure has a counterpart worth knowing about. Microsoft Learn lists Claude Mythos 5.1, Mythos 5 and Mythos Preview in Foundry as a gated research preview, with access granted at Anthropic's discretion and prioritised for defensive cybersecurity use cases. If your security team wants a frontier cyber model inside the Azure tenant it already governs, that is the application to file first. If it wants the one that tops CWE-bench and Wiz's pentest set today, Fairwind is the form, and the model runs on Google's infrastructure.

The Swedish and EU angle

A third supplier is a third set of terms. Claude in Foundry is already a Non-Microsoft Product under the Product Terms, with Anthropic's data terms applying. Argon would sit entirely outside the Microsoft agreement: Google Cloud terms, Google's data processing addendum, Google's sub-processor list. For a public-sector buyer that is a new supplier assessment, not an addendum to an existing one.

Residency is available, but not in Sweden. The EU multi-region endpoint keeps processing inside the Union, which satisfies most Swedish policies that say "within the EU/EES". A policy that says "within Sweden" cannot be met on Gemini today, since Stockholm is not on the generative AI locations list, and it cannot be met on Foundry's EU Data Zone either, which is EU-wide. The requirement, not the vendor, is what needs clarifying.

The AI Act applies regardless of which cloud the model sits in. The Regulation has applied in general since 2 August 2026, and the Commission has said there is no pause. The general-purpose model obligations fall on Google as provider. Your obligations as deployer, including logging, human oversight and transparency where they apply, are identical for Argon, Sol and Sonnet. A multi-cloud model estate makes the logging obligation harder to meet uniformly, which is a reason to front every model with the same gateway.

Introductory pricing and public procurement do not mix. A price that will double on an unannounced date is hard to put in a framework agreement. If Argon ends up in a Swedish tender, price it at $4/$20 and record the introductory rate as a note.

Decision guide

  • Running agents in Foundry today? Nothing changes this week. Argon has no model ID, no enterprise listing and no residency option yet. Keep the Sonnet 5.5 or GPT-6.1 Sol decision from last week.
  • Long-document finance or legal workloads? Argon's largest published leads are here. Build the eval set now so you can run it the day an API ID appears, and run it against GPT-6.1 Sol and Opus 5.5 in Foundry at the same time.
  • Terminal-heavy coding agents? The charts favour Opus 5.5 and Astra. Argon's DeepSWE lead is 3.7 points; its FrontierSWE deficit is 7 to 10. Do not switch on the headline.
  • Need EU processing? Wait for Argon on the EU multi-region endpoint of Gemini Enterprise Agent Platform. The Developer API route is for evaluation only.
  • Critical infrastructure security team? Apply to Fairwind for Argon on Google, and to Anthropic's Mythos gated preview in Foundry. Compare both inside your own environment before choosing.
  • Writing a 2027 budget? Use $4/$20. The $2/$10 rate has no published end date, and Google's last introductory price ended on a fixed day with a doubling.
  • Want to be ready without committing? Import the Gemini OpenAI-compatible endpoint into API Management now, with a token quota policy, so adding Argon later is a model-name change behind a gateway you already govern.

Three vendors at the same price in one week is good news for buyers and bad news for anyone who hard-coded a model name. The pattern that survives this pace is the one in our model exit plan: one gateway, one eval set, one usage log, and a model string that lives in configuration.

subscribe # the AI news that matters, minus the noise

Book a Call

Sources

Tags

Related posts