Google Launches Gemini 3.6 Flash and 3.5 Flash-Lite, Teasing Gemini 4 While Pro Model Lags Behind

Google Launches Gemini 3.6 Flash and 3.5 Flash-Lite, Teasing Gemini 4 While Pro Model Lags Behind

If you were waiting on Google to drop Gemini 3.5 Pro, you’ll have to keep waiting. Instead of delivering its next flagship heavy-hitter, Mountain View took a different route by beefing up its workhorse layer with three new releases: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a security-focused model called Gemini 3.5 Flash Cyber.

Rather than chasing leaderboard headlines with massive compute models, Google’s strategy here is crystal clear: win on developer economics, output speed, and token efficiency where the actual production volume lives.

Below is every single detail, benchmark jump, and availability update from Google’s announcement.

1. Gemini 3.6 Flash: Cutting Output Tokens and Boosting Execution

Following up on the initial 3.5 Flash release from I/O, Gemini 3.6 Flash responds directly to months of enterprise feedback. The headline upgrade is how much less text the model needs to generate to solve a problem.

According to data from the Artificial Analysis Index, 3.6 Flash uses roughly 17% fewer output tokens than 3.5 Flash. For agentic setups and multi-step tasks, Google points out that the model takes significantly fewer reasoning steps and tool calls to get to the finish line. On specific coding benchmarks like Datacurve’s DeepSWE, that token reduction actually reaches up to 65%, while tests with corporate partners like Harvey showed tasks getting completed about 12% faster on average.

Pricing and Core Specs

  • API Pricing: $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. That’s a noticeable drop from the $9.00 output rate on 3.5 Flash.
  • Knowledge Cutoff: Jumped over a year ahead, moving from January 2025 up to March 2026.
  • Context & Output Window: Maintains a 1-million-token context window along with a 64k output token cap.
  • Built-in Capabilities: Supports client-side computer control (OS manipulation), parallel function calling, structured data outputs, and native handling of text, images, video, audio, and PDFs.

Benchmark Improvements Over 3.5 Flash

  • DeepSWE (Software Engineering): Rose to 49% (up from 37%), delivering precision fixes with fewer execution loops and unwanted edits.
  • MLE Bench (Machine Learning Research): Jumped to 63.9% (up from 49.7%).
  • OSWorld-Verified (Computer Use): Hit 83.0% (up from 78.4%).
  • GDPval-AA (Knowledge Work): Reached a score of 1421 (up from 1349).

2. Gemini 3.5 Flash-Lite: The High-Speed Volume Engine

For developers handling mass data pipelines, agentic search, or real-time document processing, Google introduced Gemini 3.5 Flash-Lite. Positioned as the fastest and cheapest option in the 3.5 tier, it pumps out up to 350 output tokens per second.

It replaces the older 3.1 Flash-Lite from March with massive generational leaps:

  • Terminal-Bench 2.1 (Agentic Tasks): Jumped to 54% (up from 31%).
  • GDM-MRCR v2 (Long Context Retrieval): Scored 72.2% (up from 60.1%).
  • GDPval-AA v2: Rose sharply to 1140 (nearly doubling from 642).

In fact, Google claims 3.5 Flash-Lite comfortably beats the original base Gemini 3 model across several major metrics, including SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%).

Pricing & Thinking Controls

  • Pricing: Unbeatably low at $0.30 per 1M input tokens and $2.50 per 1M output tokens.
  • Configurable Reasoning: Includes adjustable “thinking” levels (minimal, low, and high). Developers can dial down reasoning to prioritize sub-second latency on basic lookups or turn it up to give sub-agents headroom for complex jobs.

3. Gemini 3.5 Flash Cyber & CodeMender

Rounding out the release is Gemini 3.5 Flash Cyber, a specialized model tuned exclusively to find, validate, and patch software vulnerabilities before exploits happen in the wild.

Because automated vulnerability scanning carries dual-use risks (potential misuse for cyber-offense), Google is holding back on a self-serve API launch. Instead, it’s being deployed inside Google’s CodeMender multi-agent tool as part of a limited-access pilot program reserved for government bodies and trusted enterprise security partners.

What About Gemini 3.5 Pro and Gemini 4?

Conspicuously missing from the event was Gemini 3.5 Pro, which has now missed its expected launch timeline multiple times.

Google acknowledged the delay, stating simply that 3.5 Pro remains in private partner testing and will launch broadly “as soon as it’s ready.” However, the DeepMind team confirmed they’ve already pivoted heavy resources forward, officially kicking off their largest pre-training run to date for Gemini 4.

Where to Find the New Releases Today

  • Developers: Both Gemini 3.6 Flash and 3.5 Flash-Lite are live in the Gemini API through Google AI Studio, Android Studio, and Google Antigravity.
  • Enterprises: Available through the Gemini Enterprise Agent Platform, with 3.6 Flash also accessible via the Gemini Enterprise app.
  • Everyday Users: Both models are rolling out inside the main Gemini App, while 3.5 Flash-Lite is being integrated directly into Google Search to speed up overview responses.
Author
Related Posts