Skip to main content
TechnologyJul 22, 2026· 4 min read

Google Launches Gemini 3.6 Flash and New 3.5 Models: Can 350 Tokens Per Second Be Enough?

Google Launches Gemini 3.6 Flash and New 3.5 Models: Can 350 Tokens Per Second Be Enough?

Google has added three new models to its Gemini family.

Gemini 3.6 Flash

Gemini 3.6 Flash replaces 3.5 Flash as the flagship model of the series, while 3.5 Flash-Lite becomes the fastest and most economical variant of the 3.5 generation. Additionally, 3.5 Flash Cyber is a model dedicated to researching and correcting cyber vulnerabilities, integrated within the CodeMender agent. The common theme behind these three announcements is to make the execution of large-scale AI agents more sustainable in terms of costs and latency.

Gemini 3.6 Flash is based on feedback collected from developers and companies already using 3.5 Flash. The main leap concerns efficiency: according to the Artificial Analysis Index, the new model consumes 17% fewer output tokens than its predecessor, with peaks reaching up to 65% on benchmarks like Datacurve's DeepSWE. This translates to a lower price as well: $1.50 per million input tokens and $7.50 per million output tokens, compared to previous listings.

On the performance front, the numbers indicate a model that is not only more accurate but also more economical. On DeepSWE, the score increases to 49% from 37% of 3.5 Flash, while on MLE Bench, dedicated to machine learning research, the leap is from 49.7% to 63.9%.

The computer control capabilities also improve, achieving an 83.0% score on OSWorld-Verified, up from the previous 78.4%, and the computer use function becomes a native client-side tool for both the Gemini API and Gemini Enterprise. In terms of knowledge work, the GDPval-AA v2 score rises from 1349 to 1421, a figure that companies like Hebbia and Harvey have already found useful for document, graph, and report analysis.

Google has also integrated enhanced security safeguards in 3.6 Flash under Frontier Safety, related to chemical, biological, radiological, and nuclear (CBRN) risks and offensive cyber use. The model is more resistant to jailbreak attempts, even though it has been trained to reduce waste in legitimate use cases.

3.5 Flash-Lite: Speed as a Top Priority

Alongside the flagship model comes Gemini 3.5 Flash-Lite, designed for low-latency activities and high-throughput workloads such as agent search and document processing. It is the fastest model in the entire 3.5 series, with 350 output tokens per second according to Artificial Analysis measurements, at a price of $0.3 per million input tokens and $2.5 per million output tokens.

Compared to 3.1 Flash-Lite, the qualitative improvement is evident on several fronts: on Terminal-Bench 2.1, the score rises from 31% to 54%, while on long context management (GDM-MRCR v2), it jumps from 60.1% to 72.2%. The leap is even more pronounced on GDPval-AA v2, where the score nearly doubles from 642 to 1140. The most surprising data is that 3.5 Flash-Lite even surpasses the 3 Flash model on SWE-Bench Pro (54.2% vs. 49.6%) and on OSWorld-Verified (74.0% vs. 65.1%), while still being an economical variant. Here too, the computer use function is integrated as standard, and developers can adjust the reasoning level of the model, from the minimum designed for high-volume, low-latency tasks to higher levels for more complex multi-step workloads.

The Third Element of the Announcement: Gemini 3.5 Flash Cyber

The third piece of the announcement is Gemini 3.5 Flash Cyber, a model derived from 3.5 Flash but specifically refined to identify and correct vulnerabilities in code. It is employed within CodeMender, Google's security agent, where multiple instances of the model work in parallel to produce a single consolidated report. On the industry benchmark, CyberGym, the system achieves competitive results with larger models but at a significantly lower cost per token. Being a dual-use technology, Google has opted for a controlled distribution: 3.5 Flash Cyber will be available only to governments and selected partners through CodeMender, as part of a limited-access pilot program coming in the next few months.

Google has also taken the opportunity to update the roadmap for upcoming models. Gemini 3.5 Pro is currently being tested with selected partners and will be made publicly available as soon as it is ready. Meanwhile, the team has already launched what is described as the most ambitious pre-training ever conducted, a prerequisite for the development of Gemini 4.

In terms of availability, both 3.6 Flash and 3.5 Flash-Lite are accessible from today. Developers can find them on Google AI Studio, Android Studio, and, limited to 3.6 Flash, on Google Antigravity. Companies can use them via the Gemini Enterprise Agent Platform, while end-users can find them in the Gemini app, with 3.5 Flash-Lite also being distributed on Google Search.