Google's New Gemini Model Could Quietly Inflate Your AI Bill

3 September 2026

Google released Gemini 3.8 Flash on 2 September 2026 — its third Flash-branded model in six weeks, according to Ars Technica and The Verge. The new model adds three 'thinking levels': low, medium and high, letting it reason more deeply on complex requests. Google itself has flagged that the harder the model thinks, the more it can cost to run.

Most small business owners never touch Gemini directly. You meet it inside tools built on top of it — a quoting chatbot on your website, a customer service auto-responder, an email triage assistant sold to you as part of a subscription. When Google upgrades the model underneath those tools, the tool often upgrades with it, sometimes automatically, and the running costs behind it can shift too.

What this means for your margin

Say your website's enquiry bot uses Gemini Flash to answer messages that land overnight. If your supplier defaults every enquiry to 'high' thinking because it produces slightly better answers, then every simple message — 'what's your day rate', 'are you free next week' — could now run at the more expensive reasoning setting, even though a one-line reply never needed deep reasoning in the first place. Multiply that across hundreds of enquiries a month and you're paying premium prices for routine questions.

The fix is not complicated. It is asking your supplier a direct question: which thinking level does our tool actually run on, and can we set it to 'low' for routine tasks and 'high' only for the complex ones — a detailed quote, a contract summary, a technical query? Most SaaS tools built on Gemini's API will have a setting for this, even if it is buried in an admin panel you have never opened.

What to do this month

Prompted by: https://www.theverge.com/ai-artificial-intelligence/988742/google-gemini-3-8-flash

Want this level of clarity on your own numbers?

Start with the free Margin vs Volume calculator — 30 seconds, no sign-up.

Run your numbers →