Skip to content
Putting technology to work.
Insights to guide decisions and action.

Search articles

Three new Gemini models debut — How SMBs should choose "which model at what cost"

Table of contents · 6 items

"We really want to adopt AI into our workflows. But model names keep multiplying one after another, and we honestly can't tell which one to use. The pricing doesn't seem uniform either, and for a company of our scale, we have no idea what our monthly costs will look like"—we received this inquiry from someone managing information systems alongside other duties at a manufacturing company. The desire to experiment is there. But the more options expand, the harder it is to take action. This is by no means an uncommon dilemma.

On July 21, 2026, Google announced three lightweight models at once. Looking purely at the names, you might feel like "another set has arrived." However, for anyone struggling with model selection, this is an excellent opportunity to establish a clear decision framework. Rather than advocating for a specific model, we use this announcement as a foundation to structure how to match the right weight of model to the right use case with a realistic sense of cost.

What was announced

The models unveiled are all part of Gemini's "Flash" family—lightweight models positioned in the high-speed, low-cost tier. Distinct from top-tier, high-precision models, they can be understood as a lineup designed to process high-volume workloads and daily tasks quickly and affordably.

ModelPositioningIntended use case
Gemini 3.6 FlashGeneral-purpose workhorse; reduced output pricing compared to previous-generation 3.5 FlashCoding assistance, general knowledge work
Gemini 3.5 Flash-LiteLightest model in the 3.5 tier; faster with lower unit pricingHigh-volume batch processing, high-throughput tasks such as summarization and classification
Gemini 3.5 Flash CyberBuilt on 3.5 Flash; specialized in discovering, verifying, and patching code vulnerabilitiesSecurity use cases (restricted availability)

Each of the three models serves a distinct role. Gemini 3.6 Flash is positioned as the "primary workhorse for daily use," improved over the previous generation to complete identical tasks using fewer output tokens. Reports indicate that efficiency has improved by executing the same work with fewer output tokens, alongside a price reduction for output tokens (reported as dropping from $9.00 to $7.50 per million output tokens, with input holding at $1.50).

Gemini 3.5 Flash-Lite is the lightest model, built for "sheer speed and handling massive volumes." Being lighter and faster than 3.6 Flash with lower unit pricing, it is well-suited for automated processing of large volumes of text.

Gemini 3.5 Flash Cyber takes a different angle, specializing in detecting and patching code vulnerabilities. Operating within an agent called CodeMender, it reportedly achieves strong performance on the CyberGym benchmark; however, due to misuse concerns, access is restricted to government agencies and select partners, meaning it is not immediately available to everyday SMBs. Alongside these releases, parallel development on the next-generation "Gemini 4" was also announced.

Why are models multiplying?

The issue troubling the person in our opening scenario was precisely this endless expansion of options. However, once you understand why vendors release more models, navigating your choices becomes significantly clearer.

The reason is simple: satisfying every use case with a single model always leads to waste. While complex reasoning, such as interpreting risks in contracts, requires a high-precision model, using that same model for simply routing inquiry emails into standard categories provides excess accuracy and only inflates costs. Conversely, having a lightweight model make difficult decisions may be cheap, but it results in errors. That is why vendors line up tiers from "fast and cheap" to "smart and expensive," enabling users to choose according to their specific needs.

In other words, the increasing number of models is not about imposing options; rather, it is closer to the reality that the adjustment dials for trade-offs are becoming more granular. Instead of being intimidated by the increased variety, it is far more productive to consider where your work belongs on this dial. As a foundation for this perspective, our previously compiled guide to Gemini API cost tiers is also helpful.

Choosing models by use case

So how should you actually map your company's operations? The key here is not choosing based on brand names or the novelty of version numbers. The criteria boil down to two points: how much intelligence the task requires, and the volume of work you need to process.

For high-volume, simple, routine tasks, head straight for lightweight models without hesitation. Categorizing inquiry emails, summarizing articles and meeting minutes, generating standardized product descriptions, tagging open-ended survey responses—lightest-weight models like Flash-Lite are ideal for these jobs where decisions are simple but transaction volumes are high. The accuracy difference per item is small, while processing speed and unit costs make a much greater impact.

For everyday knowledge work, rely on balanced flagship models. Drafting emails, outlining internal documents, simple coding assistance, primary organization for research—tasks that require some thought without demanding specialized judgment fall within the domain of flagship models like 3.6 Flash. Striking a balance between affordability and intelligence, this tier is more than sufficient to handle the daily operations of most SMBs.

For complex reasoning and heavy decision-making, do not hesitate to use top-tier models. Comparing contract clauses, highly specialized analysis, and design decisions requiring multi-step logic will encounter oversights with lightweight models. If you downgrade to a lighter model here simply for cost reasons, the expense of rework and errors will ultimately outweigh the savings. The fewer the items and the higher the stakes per case, the more value there is in entrusting them to top-tier models.

The trick is to follow the sequence of "testing with a lighter model first, and only upgrading tasks whose accuracy falls short." Running everything on high-precision models from the start will cause costs to skyrocket. Even with the same lightweight model, practical accuracy can often be achieved depending on how prompts are phrased, making it well worth testing before jumping straight to a top-tier model.

How to think about costs

Simply staring at numbers on a pricing table makes it difficult to visualize your company's actual financial burden. To understand costs, you need to look from a slightly different angle.

Pay-as-you-go generative AI billing is broadly divided into input tokens (the volume of text read by the AI) and output tokens (the volume of text returned by the AI), with output tokens typically set at a higher unit price. When viewed from this perspective, the significance of 3.6 Flash improving to handle "the same task with fewer output tokens" while also lowering output unit prices becomes crystal clear. You can complete the same task with fewer high-priced tokens at a lower unit rate—meaning you cannot grasp true costs without looking not just at a single line on the pricing table, but at how many tokens are actually consumed per request.

For SMBs estimating costs, the realistic procedure consists of the following three steps. First, narrow down to a single workflow you want to delegate to AI, and roughly measure the input and output volume per transaction. Second, multiply that by monthly volume to calculate total token usage. Third, apply the unit prices of candidate models to estimate monthly costs. Following this sequence allows you to make decisions like "starting small with a lightweight model, and expanding scope once results are demonstrated." Jumping straight into total costs for a company-wide rollout almost always leads to hesitation from intimidating numbers.

Another factor that cannot be overlooked is that, separate from pay-as-you-go API billing, there is the option of using Gemini as part of subscriptions like Google Workspace. Depending on your usage patterns, this can make cost management much simpler. Which approach fits your company depends on your operational style, and we have gathered decision criteria in our guide to choosing Gemini plans on Workspace.

Step-by-step rollout sequence for your company

Finally, let us summarize the steps for taking action. The secret to avoiding missteps is to keep your company's actual workflows at the center, rather than debating a migration every time a new model is released.

The way to start is simple: first, select just one high-volume, simple task that you currently most want AI to take off your hands. Next, test it with a lightweight model, and if accuracy is insufficient, move up one tier. Running this small loop for a single workflow to grasp cost-efficiency and effectiveness before expanding to the next ensures you will not be swayed no matter how many models hit the market.

However, the biggest stumbling block here is not technical, but rather "implementing it only for it to go unused." Even a carefully chosen model is meaningless if nobody on the ground touches it. How to avoid this pitfall is covered in detail in our article on the issue of Gemini going unused after adoption.

There are two things to do first. One is to list out a single high-volume, simple task within your company. The other is to test it on a small scale with a lightweight model, confirming the cost per transaction and accuracy with your own eyes. Starting from here will ensure that no matter what model comes out next, you can evaluate it calmly using your own yardstick.

If you would like to discuss specifically which model to apply to your company's workflows and how to configure it, including cost considerations, please feel free to reach out through our consultation on development, AI, and automation. We will review your situation and provide tailored recommendations.

Sources

Share this articleXFacebook
Kakeru Suzuki

Fascinated by the possibilities of technology, has had a deep interest in programming and digital art since student days

Turn this article's theme into your company's next step

Thinking together, starting from the work you entrust to AI.

We organize your current operations and data to define the scope entrusted to AI, what humans should review, and how to run trials.

  • Target operations
  • Data to use
  • How to verify effectiveness
Consult on AI adoption for your business

You can consult with us from the initial conceptual stage. Details from this article will be carried over to the inquiry form.

Receive the latest articles by email