Gemini 3.5 Flash Is Here. The Pro Model Still Isn’t.
On this page
Google has released Gemini 3.5, its latest model family, and the pitch is in the tagline: frontier intelligence with action. The series kicks off with 3.5 Flash, which Google says is available to billions of people globally. The flagship Gemini 3.5 Pro, announced alongside it, remains listed as “coming soon” on Google’s own model page. Here is what actually shipped.

What Gemini 3.5 Flash Does
Google positions 3.5 Flash around agents rather than chat. The capabilities it leads with are agentic coding, advanced multimodal understanding, long-horizon tasks meant to run over extended timeframes, and multi-step problem solving with tools.
That framing shows up in the demos: generating six payment UI options in under a minute, ingesting the AlphaGo paper and building a working game autonomously, coordinating subagents to design a virtual city, deploying parallel agents to rename and restructure messy datasets.
Where you can use it today:
-
Everyone, through the Gemini app and AI Mode in Google Search
-
Developers, via Google Antigravity, the Gemini API in Google AI Studio, and Android Studio
-
Enterprises, through the Gemini Enterprise Agent Platform
The Gemini Benchmarks Google Published

The numbers are Google’s own, and they tell a more interesting story than a clean sweep. Flash leads on agentic and multimodal work while trailing on raw reasoning.
|
Benchmark |
Gemini 3.5 Flash |
Gemini 3.1 Pro |
Claude Opus 4.7 |
GPT-5.5 |
|
MCP Atlas (multi-step workflows) |
83.6% |
78.2% |
79.1% |
75.3% |
|
Toolathlon (real-world tool use) |
56.5% |
n/a |
n/a |
55.6% |
|
Finance Agent v2 |
57.9% |
43.0% |
51.5% |
51.8% |
|
MMMU-Pro (multimodal) |
83.6% |
80.5% |
75.2% |
81.2% |
|
CharXiv Reasoning (charts) |
84.2% |
83.3% |
82.1% |
84.1% |
|
Terminal-bench 2.1 (agentic coding) |
76.2% |
70.3% |
66.1% |
78.2% |
|
SWE-Bench Pro |
55.1% |
54.2% |
64.3% |
58.6% |
|
Humanity’s Last Exam |
40.2% |
44.4% |
46.9% |
41.4% |
|
ARC-AGI-2 |
72.1% |
77.1% |
75.8% |
84.6% |
Read the pattern rather than the highlights. A Flash-tier model beating Gemini 3.1 Pro on tool use and finance tasks is the actual news. But on Humanity’s Last Exam and ARC-AGI-2, 3.5 Flash sits below Google’s own 3.1 Pro, and well below GPT-5.5. This is a model tuned for doing things, not for the hardest abstract reasoning. Google’s positioning is honest about that.
What Early Users Say
The partner quotes Google published point the same direction. JetBrains reports coding and reasoning quality close to Gemini Pro while keeping Flash’s speed and cost profile, with low-reasoning coding performance up 10 to 20 percent over the previous Flash generation. Box measured a 19.6 percent gain over Gemini 3 Flash on its enterprise evaluation set. Security firm Armadin cites 42 percent better performance on a long-range cyber benchmark alongside a 68 percent improvement in token efficiency.
Shopify, Macquarie Bank, Salesforce, Ramp, Xero, and Databricks are all running it in production or pilots, mostly for the same job: parallel subagents chewing through long, document-heavy workflows.
The Gemini 3.5 Pro Question
Google said 3.5 Pro was already in internal use and would roll out the following month. That has not happened. The model page still carries the “3.5 Pro coming soon” label, and the Pro tier on offer is Gemini 3.1 Pro.
Plenty of specifics are circulating about what 3.5 Pro will bring, including a 2-million-token context window and premium pricing. None of it comes from Google. Until there is a model card, API documentation, or a pricing page, those are reports, not facts.
The Takeaway
Gemini 3.5 Flash is a real release doing a specific thing well: agentic execution at Flash speed and cost. For most people that is the model they will actually touch, through the Gemini app or AI Mode, at no charge.
The flagship is a different matter. Google announced 3.5 Pro, said it was coming, and has not shipped it, while OpenAI and others have moved. Whenever it lands, the number worth watching is not the size of the context window but whether reasoning quality holds across it. We will cover it when there is something official to cover.
❓ Frequently Asked Questions
Answers to relevant questions about this AI tool