AI & Machine Learning

Gemini 3.5 Flash vs GPT-5.5 vs Claude Opus 4.7 at I/O 2026

By Technspire TeamMay 20, 20267 views

Google opened I/O 2026 yesterday at Shoreline Amphitheatre in Mountain View. The conference runs 19-20 May, and the day-one keynote declared the start of what Google calls the "agentic Gemini era". The headline for engineering teams: Gemini 3.5 Flash is generally available from day one, pitched as frontier intelligence for agents and coding, with Gemini 3.5 Pro following next month. Alongside it came Gemini Omni, a model that generates video from any mix of text, image, audio and video input. Google also introduced a new $100/month Google AI Ultra plan aimed squarely at developers, plus a substantial upgrade to its Antigravity agent platform. If your organisation is standardised on Azure, as most of our Swedish enterprise clients are, this is not a keynote you can file under "interesting, but not our cloud". Frontier releases from any of the big three labs reset the price/performance baseline that every vendor gets negotiated against, including the ones already in your tenant. Three frontier-class agentic coding models have now shipped in five weeks: Claude Opus 4.7 on 16 April, GPT-5.5 on 23 April, and Gemini 3.5 Flash yesterday. That changes the calculus, even for teams that never leave Azure.

What Google actually announced

Gemini 3.5 Flash now, Gemini 3.5 Pro next month

The naming order is deliberate. Google is leading the 3.5 generation with Flash, historically its mid-tier, price-efficient line, and positioning it as a frontier agentic and coding model in its own right, with what Google describes as powerful new action-taking capabilities designed for complex, multi-step agentic workflows. Gemini 3.5 Flash is generally available immediately through the Gemini API, Google AI Studio, Android Studio and Google Antigravity. Gemini 3.5 Pro is still in internal testing, with a rollout planned for next month. Leading with the cheaper model is a statement of intent. Google wants the volume workloads where token economics dominate the decision: agent loops, coding assistants, background automation.

Antigravity 2.0 and managed agents

Google Antigravity, the company's agentic development environment, moves to version 2.0 as a standalone desktop application with multi-agent orchestration, joined by a new Antigravity CLI for terminal-based agent workflows. On the infrastructure side, the Gemini API gains Managed Agents: provisioned agents that run in sandboxed Linux environments on Google's side, rather than on infrastructure you operate. That is a direct play for the same territory Azure AI Foundry's agent tooling occupies. The pitch: you bring a task definition, and Google runs everything else, loop and sandbox included.

A $100/month developer tier

Google introduced a Google AI Ultra plan for developers at $100 per month with five times higher usage limits. Per-seat subscription pricing for AI coding tools has been drifting upward across the industry, and Google is now anchoring a number in the middle of the range and attaching its newest model to it on day one. Whether it is cheap or expensive for your team depends entirely on usage patterns. It is, however, a concrete figure your procurement team will quote in the next Copilot or Claude renewal discussion, which is precisely why Google published it.

Gemini Omni and agents on Android

Gemini Omni generates and edits video from mixed inputs (text, images, audio, existing video), and every output carries an imperceptible SynthID watermark. We will come back to why that watermark matters more in the EU than anywhere else. On the consumer side, Android gains Halo, a dedicated space on the phone for monitoring what agents are doing in the background. The feature is small; the implication is not. Google expects agents to be running unattended on end-user devices as a normal state of affairs.

Three frontier agentic models in five weeks

Step back from the keynote and look at the calendar. On 16 April, Anthropic released Claude Opus 4.7: improved long-horizon agentic coding, higher-resolution vision input up to 2,576 pixels on the long edge, a new xhigh effort level, priced at $5 per million input tokens and $25 per million output tokens, and available from day one on the Claude API, Amazon Bedrock, Google Vertex AI and, importantly for this audience, Microsoft Foundry. On 23 April, OpenAI announced GPT-5.5, with API availability the following day and improvements in agentic coding, reasoning and knowledge work among the headline claims. And yesterday, Gemini 3.5 Flash went GA. Every one of these releases leads its marketing with the same two words: agentic and coding. The labs have converged on the same battleground because that is where enterprise spend is concentrating, on models that can plan, call tools, edit repositories and run for minutes or hours without a human in the loop.

A note on benchmarks: Gemini 3.5 Flash is one day old. Vendor launch claims, Google's included, are claims, not evidence. Independent evaluations and community testing will take a week or two to accumulate, and our consistent experience across the last several model generations is that launch-day positioning and real-world agentic performance on your codebase correlate loosely at best. No benchmark ranking between these three models is asserted here, because as of this morning no trustworthy independent comparison exists.

The calculus for Azure-first teams

Where each model actually lives

For a team standardised on Azure, the platform question usually settles the model question before quality ever gets measured. OpenAI's frontier models reach you natively through Azure AI Foundry, inside your existing tenant, DPA and networking. Anthropic's Opus 4.7 is available through Microsoft Foundry as well, a distribution fact that has quietly made three-model bake-offs inside a single Azure subscription possible in a way they were not eighteen months ago. Gemini is the exception: Gemini 3.5 Flash is consumed through the Gemini API, Google AI Studio or Google Cloud, not through Azure. Adopting it means a second AI vendor relationship at minimum, and in most enterprise setups a second cloud footprint: separate identity, separate networking, separate data processing agreement, separate egress path for your code and data.

That asymmetry cuts both ways. It raises the bar Gemini must clear on merit: it has to be better enough to justify the integration and governance overhead. But it also means that if Gemini 3.5 Flash turns out to deliver frontier agentic coding at Flash-tier prices, the gap between "what we pay inside Azure" and "what the market charges" becomes visible and quantifiable. That number has a way of finding its way into renewal negotiations even when nobody intends to switch.

Price pressure works even when you do not switch

This is the part of the calculus that Azure-first teams most often undervalue. You do not need to run a single Gemini token in production for yesterday's announcements to save you money. Anthropic held Opus 4.7 at Opus 4.6 pricing in April rather than raising it. Google leads its new generation with its efficiency line and a $100 developer tier. These are moves in the same game: nobody in a three-way race can afford a price umbrella. If you have committed spend on Azure OpenAI or Foundry-hosted models coming up for renewal, the strongest input to that negotiation is a documented, current view of what the equivalent workload costs on the two alternatives. Building that view is an afternoon of engineering work with the framework below, not a migration project.

Agent platform gravity is the real lock-in

Models are increasingly swappable; agent platforms are not. Antigravity 2.0, Managed Agents with sandboxed execution and Android Halo are the scaffolding Google is building around the model, exactly as Microsoft has done with Foundry and Copilot. Our advice has not changed since the first agent frameworks appeared: keep your agent logic (prompts, tool definitions, orchestration, evals) in your own repository behind a thin provider interface, and treat vendor agent platforms as execution substrates you can re-point. Let your agent definitions live natively inside one vendor's platform, and weeks like this one become stressful instead of useful.

A practical evaluation framework for this week

The decision process we recommend to clients after a launch like this is deliberately boring and deliberately cheap:

  • 1. Decide whether you are even in the market. If your current model clears your quality bar and your spend is under control, log the announcement and re-read this list at renewal time. Chasing every frontier release is a tax, not a strategy.
  • 2. Wait for independent signal, about two weeks. Let independent evals and practitioner reports on Gemini 3.5 Flash accumulate before spending your own engineering time. Day-one vendor demos are the least informative data you will ever get about a model.
  • 3. Evaluate on your own tasks, not public benchmarks. A fixed set of 20-50 real tasks from your own backlog (bug fixes, refactors, agent tool-call sequences, document extractions), scored the same way against every candidate. Public leaderboards tell you about the leaderboard.
  • 4. Measure the full agentic cost, not the token price. Agentic workloads consume tokens in loops. A cheaper model that needs more iterations, more retries or more human correction is not cheaper. Record tokens, wall-clock time and success rate per completed task.
  • 5. Price the governance overhead in full. For a non-Azure vendor, add the one-off cost of DPA review, data-residency assessment, network integration and security review to the business case. For Swedish enterprises this is rarely under a few weeks of lead time.
  • 6. Convert the result into negotiation leverage first. Even a losing alternative that scores within 10-15% of your incumbent at a lower price is worth a page in your renewal file.

The comparison harness does not need to be sophisticated. A provider-agnostic task runner of thirty lines gets you most of the way:

// eval-run.mjs — same tasks, three providers, one CSV
const providers = [
  { name: 'azure-gpt',   run: (task) => callAzureFoundry('gpt-5.5', task) },
  { name: 'azure-opus',  run: (task) => callAzureFoundry('claude-opus-4-7', task) },
  { name: 'gemini',      run: (task) => callGeminiApi('gemini-3.5-flash', task) },
];

const tasks = await loadTasks('./eval-tasks/');   // your real backlog items

for (const task of tasks) {
  for (const p of providers) {
    const t0 = Date.now();
    const result = await p.run(task);
    results.push({
      task: task.id,
      provider: p.name,
      passed: await task.check(result.output),    // deterministic check per task
      inputTokens: result.usage.input,
      outputTokens: result.usage.output,
      iterations: result.iterations,              // agent loop count
      ms: Date.now() - t0,
    });
  }
}
await writeCsv('results.csv', results);

The discipline is in the task set and the deterministic checks, not the runner. Build those once and every future model launch costs you an afternoon to evaluate instead of a meeting series to debate. On current form, there will be another one within weeks.

The Swedish and EU angle

For Swedish enterprises, three EU-specific considerations sit on top of the general calculus.

Data residency and processing agreements. Most of our clients have spent the past two years consolidating AI workloads into their Microsoft estate precisely because the data-protection posture (EU data boundary commitments, existing DPAs, established sub-processor lists) was already negotiated. Bringing Gemini into scope means running that assessment again for Google: where are prompts and outputs processed, what are the retention and training-use terms for the API tier you are on, and does your sector regulator (Finansinspektionen, IMY guidance for public sector) have a view? None of this is prohibitive; Google Cloud has EU regions and enterprise terms. But it is real lead time that belongs in the business case, not discovered after a pilot succeeds.

The AI Act transparency clock. Gemini Omni watermarking every generated video with SynthID is not just a trust-and-safety gesture. The EU AI Act's transparency obligations for synthetic content, machine-readable marking of AI-generated audio, image and video, apply from August 2026, which is now under three months away. If generative video is anywhere on your roadmap, provider-side watermarking that ships by default is a genuine compliance feature. Ask every vendor in your stack, not just Google, what their equivalent answer is before that deadline.

Procurement leverage is an EU sport. Swedish enterprise and public-sector procurement runs on documented alternatives. A current, evidence-based comparison of the three frontier vendors strengthens your position under both framework agreements and direct negotiations, even when it concludes "stay on Azure". The five-week cadence of releases from Anthropic, OpenAI and now Google means that comparison goes stale quickly. Maintained continuously rather than rebuilt at each renewal, it consistently produces better terms.

What we would do this week

Do not migrate anything. Do not ignore this either. Log the concrete facts: Gemini 3.5 Flash GA yesterday, 3.5 Pro next month, a $100/month developer tier, Opus 4.7 and GPT-5.5 both under six weeks old. Then book two hours in the first week of June, once independent evaluations exist, to run your own task set against all three. If you do not have a reusable eval harness for your agentic workloads yet, that is the actual action item from Google I/O 2026. The model-choice question now recurs every few weeks, and a team that can answer it with data in an afternoon turns this release pace into negotiation leverage rather than noise. If you want help standing up that evaluation loop on Azure, including Foundry-hosted access to more than one of these model families, that is exactly the kind of engagement we do.

Sources