WritingGoogle (DeepMind / Gemini)Google (DeepMind / Gemini)published Jun 17, 2025seen 6d

Gemini 2.5: Updates to our family of thinking models

Open original ↗

Captured source

source ↗

Gemini 2.5: Updates to our family of thinking models

  • Google Developers Blog

Gemini

Gemini 2.5: Updates to our family of thinking models

JUNE 17, 2025

Shrestha Basu Mallick

Product

Google DeepMind

Logan Kilpatrick

Group Product Manager

Share

Facebook

Twitter

LinkedIn

Mail

Today we are excited to share updates across the board to our Gemini 2.5 model family:

Gemini 2.5 Pro is generally available and stable (no changes from the 06-05 preview)

Gemini 2.5 Flash is generally available and stable (no changes from the 05-20 preview, see pricing updates below)

Gemini 2.5 Flash-Lite is now available in preview

Gemini 2.5 models are thinking models, capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. Each model has control over the thinking budget, giving developers the ability to choose when and how much the model “thinks” before generating a response.

Overview of our family of Gemini 2.5 thinking models

Introducing Gemini 2.5 Flash-Lite Today, we’re introducing 2.5 Flash-Lite in preview with the lowest latency and cost in the 2.5 model family. It’s designed as a cost-effective upgrade from our previous 1.5 and 2.0 Flash models. It also offers better performance across most evals, and lower time to first token while also achieving higher tokens per second decode. This model is great for high throughput tasks like classification or summarization at scale. Gemini 2.5 Flash-Lite is a reasoning model, which allows for dynamic control of the thinking budget with an API parameter. Because Flash-Lite is optimized for cost and speed, “thinking” is off by default, unlike our other models. 2.5 Flash-Lite also supports all of our native tools like Grounding with Google Search, Code Execution, and URL Context in addition to function calling.

Benchmarks for Gemini 2.5 Flash-Lite

Updates to Gemini 2.5 Flash and pricing Over the last year, our research teams have continued to push the pareto frontier with our Flash model series. When 2.5 Flash was initially announced, we had not yet finalized the capabilities for 2.5 Flash-Lite. We also launched with a “thinking” and “non-thinking price”, which led to developer confusion.

With the stable version of Gemini 2.5 Flash rolling out (which is the same 05-20 model preview we made available at Google I/O), and the incredible performance of 2.5 Flash, we are updating the pricing for 2.5 Flash:

$0.30 / 1M input tokens (*up from $0.15 input)

$2.50 / 1M output tokens (*down from $3.50 output)

We removed the thinking vs. non-thinking price difference

We kept a single price tier regardless of input token size

While we strive to maintain consistent pricing between preview and stable releases to minimize disruption, this is a specific adjustment reflecting Flash’s exceptional value, still offering the best cost-per-intelligence available. And with Gemini 2.5 Flash-Lite, we now have an even lower cost option (with or without thinking) for cost and latency sensitive use cases that require less model intelligence.

Pricing updates for our Gemini Flash family

If you are using the Gemini 2.5 Flash Preview 04-17 , the existing preview pricing will remain in effect until its planned deprecation on July 15, 2025, at which point that model endpoint will be turned off. You can transition to the generally available model “gemini-2.5-flash”, or switch to 2.5 Flash-Lite Preview as a lower cost option.

Continued growth of Gemini 2.5 Pro The growth and demand for Gemini 2.5 Pro continues to be the steepest of any of our models we have ever seen. To allow more customers to build on this model in production, we are making the 06-05 version of the model stable, with the same pareto frontier price point as before. We expect that cases where you need the highest intelligence and most capabilities are where you will see Pro shine, like coding and agentic tasks. Gemini 2.5 Pro is at the heart of many of the most loved developer tools.

Top developer tools using Gemini 2.5 Pro

If you are using 2.5 Pro Preview 05-06, the model will remain available until June 19, 2025 and then will be turned off. If you are using 2.5 Pro Preview 06-05, you can simply update your model string to “gemini-2.5-pro”. We can’t wait to see even more domains benefit from the intelligence of 2.5 Pro and look forward to sharing more about scaling beyond Pro in the near future.

posted in:

Gemini

AI

Announcements

Gemini 2.5 Flash-Lite

AI models

Gemini 2.5 Pro

Vertex AI

Developer Tools

Gemini 2.5

Gemini 2.5 Flash

Machine Learning

Google AI Studio

Previous

Next

Related Posts

AI

Announcements

Explore

Gemma 4 12B: The Developer Guide

JUNE 3, 2026

Gemini

Web

AI

Tutorials

How-To Guides

Turn creative prompts into interactive XR experiences with Gemini

FEB. 19, 2026

AI

Cloud

Announcements

Learn

Introducing the Google Colab CLI

JUNE 5, 2026

Gemini

Google AI Studio

AI

Events

How we built the Google I/O 2026 Save the Date experience

MARCH 3, 2026

Notability

notability 9.0/10

Major model update from Google DeepMind