NexBDM Blog
AI Token Cost: four dates on the vendor pricing pages that change what your automation costs this quarter
By NexBDM Team · 2026-09-21
AI token cost is the running price of the model under an automation, billed per million tokens. On 21 September 2026 the two vendor pricing pages most small-business automations run on carry end dates on their current rates: Google's Gemini 3.x Flash rate doubles on 1 January 2027, and OpenAI's promotional rate holds at least until 21 November 2026.
AI token cost is the running price of the model under an automation, billed per million tokens read and written. On 21 September 2026 both major vendor pricing pages carry end dates on their rates: Google doubles its Gemini 3.x Flash rate on 1 January 2027, and OpenAI's promotional GPT-5.6 Sol rate holds at least to 21 November 2026.
This is a news post, and the news is on the pricing pages themselves rather than in a press release. Both pages were read in full this morning. Google's page carries a "last updated" stamp of 16 September 2026, and its three current Flash models now print two prices each: one "through December 31, 2026" and one "starting January 1, 2027". OpenAI's page prints one sentence about a promotional rate with a floor date and no ceiling. Between them sit two more dated changes that touch the same budgets: WhatsApp service messages become billable on 1 October, which we covered from Meta's own documentation on 2 September, and Google switches off an image model on 2 October. Four dates, three vendors, one quarter. This post sets them out in order, shows what a doubling does to a budget that was set on the introductory rate, and then covers the part that matters more than any rate: how to make the meter matter less.
Key takeaways
- Google's Gemini 3.8, 3.7 and 3.6 Flash are priced at 0.75 US dollars per million input tokens and 3.75 per million output tokens through 31 December 2026, then 1.50 and 7.50 from 1 January 2027. The context caching rate and the cache storage rate double on the same day. Read on Google's pricing page, 21 September 2026.
- OpenAI's page states that "GPT-5.6 Sols promotional pricing is available at least through November 21, 2026." The current standard rate is 4.00 US dollars per million input tokens and 20.00 per million output. The page does not print what the rate becomes afterwards.
- Gemini 2.5 Flash Image is "deprecated and will be shut down on October 2, 2026". Anything that generates images through it stops working that day unless it is moved.
- WhatsApp service messages, the free-form replies a team types inside the 24 hour window, become billable per message on 1 October 2026, with no volume tiers. Nothing on Meta's main pricing page has changed since we read it on 2 September.
- The two cheapest levers on both pages are the same: cached input is billed at a tenth of the fresh input rate, and batch processing is billed at half. Neither needs a new tool. Both need the work to be shaped so it can use them.
What is an AI token cost, and why does the date matter more than the rate?
A token is the unit a language model reads and writes in, roughly three quarters of a word in English. Every call to a model is billed on two meters: tokens in (your instruction, the document, the conversation so far) and tokens out (the answer). Vendors publish the two rates per million tokens, in US dollars, and most also publish a cheaper rate for input the model has already seen recently (cached input) and a cheaper rate for work you are willing to wait for (batch). Our July guide to what AI automation really costs sets out that structure and the per-message and platform layers that sit on top of it.
What has changed since July is not the structure. It is that the rates now have dates attached. In July the story on the pricing pages was direction: mainstream rates had fallen steeply over the preceding year, and budget models had reached cents per million. That is still on the page. Gemini 2.5 Flash-Lite is still listed at 0.10 US dollars per million input tokens and 0.40 per million output. But the newest models on both pages are now priced as introductions, with a day on which the introduction ends. A budget set on the rate you see today is a budget set on a rate the vendor has already told you is temporary.
That is why the date matters more than the rate. A rate you can compare. A date you have to plan for.
The four dates, in order
1 October 2026: WhatsApp service messages become billable
This is the one that hits first and it is not a token cost, but it lands in the same budget line. From 1 October, Meta charges per message for service messages, the free-form replies a person or a bot sends inside an open 24 hour customer service window, and for utility templates sent inside that window. The service category has no volume tiers, so sending more does not earn a lower rate. We read the change off Meta's own documentation on 2 September and set it out in full in the post on what becomes billable on 1 October, with the rate derivation for South Africa in the earlier post on service messages. Meta's main pricing page was re-read this morning and has not moved since: South Africa is still listed under "Rest of Africa" with calling code 27, and the per-country rates are still served through an interactive page rather than printed inline.
2 October 2026: Google switches off Gemini 2.5 Flash Image
Google's pricing page carries a warning box on the Gemini 2.5 Flash Image row, the model it markets as Nano Banana: "deprecated and will be shut down on October 2, 2026; migrate to Gemini 3.1 Flash Image or Gemini 3.1 Flash Lite Image to avoid service disruption." If you have a workflow that generates product images, social cards or thumbnails through that model, it stops returning images in eleven days. The replacement it names, Gemini 3.1 Flash Image, is priced at 0.50 US dollars per million input tokens for text and image, 3.00 per million output tokens for text, and 60.00 per million output tokens for images, which the page itself translates to 0.067 US dollars per 1K image. A shutdown is the one kind of pricing change you cannot absorb in a budget line. It has to be done.
At least 21 November 2026: OpenAI's promotional rate on GPT-5.6 Sol
OpenAI's page prints one sentence on this, and it is worth reading exactly: "GPT-5.6 Sols promotional pricing is available at least through November 21, 2026." The standard table beside it lists gpt-5.6-sol at 4.00 US dollars per million input tokens, 0.40 per million cached input tokens, and 20.00 per million output tokens, for prompts within the short context band. Batch is exactly half: 2.00 in, 10.00 out. Flex, the wait-for-it tier, is the same half rate.
Two things the sentence does not say. It does not say what the rate becomes after the promotion, and it does not say the promotion ends on 21 November. "At least through" is a floor. The rate could hold into December or beyond. What the sentence does establish is that the number in the table is not the permanent number, and the vendor has said so. If your automation is on that model, the honest budget line carries the current rate with a note that says "promotional, floor date 21 November, post-promotion rate not published".
1 January 2027: Google doubles the Gemini 3.x Flash rate
This is the largest change and the one printed most plainly. Three models on Google's page, Gemini 3.8 Flash (which the page announces as "now available"), 3.7 Flash and 3.6 Flash, each show the same pair of prices. Input: "$0.75 through December 31, 2026. $1.50 starting January 1, 2027." Output, including thinking tokens: "$3.75 through December 31, 2026. $7.50 starting January 1, 2027." Context caching goes from 0.075 to 0.15 per million tokens, and the cache storage price from 0.50 to 1.00 per million tokens per hour. Every figure is in US dollars per million tokens, and every one of them doubles.
The older models on the same page do not carry a date. Gemini 3.5 Flash is listed at 1.50 in and 9.00 out with no introductory note, which is the number the 3.x Flash line is being drawn back toward. Gemini 2.5 Flash-Lite stays at 0.10 in and 0.40 out. So the picture on 1 January is not "Google got more expensive". It is that the newest Flash models stop being priced below the model they replaced.
What a doubling does to a budget set on the introductory rate
Take a plain automation: it reads a document and a short instruction, about 2,000 tokens in, and writes a short structured answer, about 500 tokens out. It runs a thousand times a month. That is two million input tokens and half a million output tokens a month. The table below prices that workload on the rates printed this morning. Every figure is in US dollars, from the vendor page, and the arithmetic is shown so you can check it.
| Model and rate | Input (2M tokens) | Output (0.5M tokens) | Month |
|---|---|---|---|
| Gemini 3.8 Flash, through 31 December 2026 | 2 x 0.75 = 1.50 | 0.5 x 3.75 = 1.88 | 3.38 |
| Gemini 3.8 Flash, from 1 January 2027 | 2 x 1.50 = 3.00 | 0.5 x 7.50 = 3.75 | 6.75 |
| Gemini 2.5 Flash-Lite, no date on the page | 2 x 0.10 = 0.20 | 0.5 x 0.40 = 0.20 | 0.40 |
| GPT-5.6 Sol, standard, promotional | 2 x 4.00 = 8.00 | 0.5 x 20.00 = 10.00 | 18.00 |
| GPT-5.6 Sol, batch, promotional | 2 x 2.00 = 4.00 | 0.5 x 10.00 = 5.00 | 9.00 |
At this scale the absolute numbers are small, and that is the point most owners take from a table like this: the model is not where the money goes. That is true for a thousand runs. It stops being true when the automation reads a whole inbox, a whole contract or a whole call transcript per run, because the input meter is the one that scales with the size of what you feed it. A 40,000-token contract costs twenty times what the 2,000-token document does, on every rate in the table, and the doubling applies to all of it.
The more useful reading of the table is the spread. The same workload costs 0.40 on the cheapest listed model and 18.00 on the flagship at standard rate, a factor of forty-five, and the doubling in January moves one row by 3.37. The choice of model and the shape of the work move the bill by far more than any single vendor price change does. That is where the control is.
How to make the meter matter less
Every one of these is a change to how the work is shaped, not a change of vendor. Most cost nothing to do and all of them are on the pricing pages already.
- Pin the model by name, and write its rate and its date beside it. "Gemini 3.8 Flash, 0.75 in, 3.75 out, through 31 December 2026." An automation that calls "the latest Flash model" gets whatever the alias points to, at whatever that costs on the day. OpenAI's page says this outright about its own aliases: they "will be updated to point to the latest models, with pricing adjusted to match each underlying model."
- Set the budget on the post-date rate now. If the workload in the table is yours, budget 6.75 for it, not 3.38, and treat the three months of the lower rate as the margin, not the plan.
- Cache the part of the prompt that never changes. Both pages bill cached input at a tenth of the fresh rate: 0.40 against 4.00 on GPT-5.6 Sol, 0.075 against 0.75 on Gemini 3.8 Flash. Most automations send the same long instruction, the same examples and the same reference document on every run. Structure the call so that fixed block goes first and is cached, and only the new record follows it.
- Batch the work that is not waiting for a person. Overnight summaries, end-of-day classification, monthly reports: both pages bill batch at half the standard rate. A reply to a customer cannot wait. A report that runs at 06:00 for a 08:00 meeting can.
- Capture once, so the model reads the record and not the inbox. The largest input bills come from feeding a model the same raw material repeatedly: the whole email thread, the whole document, the whole history, every time. If the enquiry, the quote and the invoice live in one record, the model reads a few hundred tokens of structured fields instead of a few thousand of prose. That is the same principle as the automation map: capture once, reuse everywhere.
- Put a rule in front of the model. A large share of the calls a chatbot or a classifier makes are answerable by a lookup: an order number, a business hour, a yes or no from a field. A rule costs nothing per run; a model call always costs something. Route to the model only what a rule cannot answer.
- Use the cheapest model that passes your test, and keep the test. The table shows a forty-five-fold spread for identical work. The only way to pick the right row is to have twenty real examples with known correct answers on disk and run every candidate model against them. A model that passes on Flash-Lite does not need Sol. Projects that fail usually never wrote the test.
- Read the pricing page on a date, not on a feeling. Put the four dates in this post in the calendar with a link to the page. On 1 October, 2 October, 21 November and 1 January, open the page, compare it with the rate written beside the model name, and change the line if it moved. It takes five minutes and it is the whole of "monitoring vendor pricing".
None of this names a NexBDM price, because none of it is one. It is what the vendors' own pages say, arranged so that the cheap levers are visible. The work of shaping an automation so that it uses cached input, batch windows and a single record is ordinary process work, and it is the reason two businesses running the same model on the same volume can have bills a long way apart.
What this changes about choosing a vendor
Less than it looks. The vendor checklist already asks who owns the model choice and who pays when the vendor's rate moves, and those two questions are the whole of this post from the buyer's side. If a proposal you are reading quotes a monthly running cost, ask three things: which model by name, which rate and which date from the vendor's page, and what happens to the line on the day the date arrives. A supplier who cannot answer the third is quoting the introductory rate as if it were permanent. The same question applies to a WhatsApp bot from 1 October: how many of its replies are service messages, and who is watching that count now that each one is billed.
Frequently Asked Questions
Will my AI bill go up on 1 January 2027?
If the automation runs on Gemini 3.8, 3.7 or 3.6 Flash, yes: input, output and caching rates all double on that date per Google's page. If it runs on Gemini 2.5 Flash or Flash-Lite, or on any model without a date beside its price, no change is printed. Check the model name first.
Does OpenAI's promotional rate end on 21 November 2026?
Not necessarily. The page says the rate is available "at least through" that date, which is a floor. It may run longer. What is certain is that the rate is labelled promotional and the post-promotion rate is not published, so a budget on that model should say so.
What is a token, in plain terms?
The unit a model reads and writes in, roughly three quarters of an English word. A 1,000-word document is about 1,300 tokens. Vendors bill per million tokens, with separate rates for tokens in and tokens out, and cheaper rates for cached input and batch work.
Should I move everything to the cheapest model?
Move each task to the cheapest model that passes a written test on real examples. The spread between the cheapest and the most expensive row on the same work is around forty-five times, so the test is worth the afternoon. Some tasks need the flagship. Most do not.
Does the WhatsApp change affect a chatbot?
Yes. From 1 October 2026 every free-form reply a bot sends inside the 24 hour window is a billable service message, with no volume tiers. A bot that sends three messages where one would do triples that line. Count replies per conversation and cut the ones that carry nothing.
The short version
Four dates this quarter change what an automation costs to run, and all four are printed on the vendors' own pages: WhatsApp service messages billable on 1 October, a Google image model switched off on 2 October, an OpenAI promotional rate with a floor of 21 November, and a doubling of the Gemini 3.x Flash rate on 1 January. The rate is the smaller lever. Pinning the model, caching the fixed prompt, batching the work that can wait, and reading one record instead of a whole inbox move the bill by more than any of the four dates do. If you want to know which model each of your automations is actually on, what it reads per run, and which of the four dates touches it, that is the kind of thing a Business Autopsy maps, and a discovery call is where it starts.
Sources, read directly on 21 September 2026: Google, "Gemini Developer API pricing", ai.google.dev, page stamped "Last updated 2026-09-16 UTC" (the Gemini 3.8, 3.7 and 3.6 Flash rates through 31 December 2026 and from 1 January 2027, the context caching and storage rates, the Gemini 3.5 Flash and 2.5 Flash-Lite rates, the Gemini 3.1 Flash Image rates, the Batch API "50% cost reduction" line, and the Gemini 2.5 Flash Image deprecation notice with its 2 October 2026 shutdown date); OpenAI, "Pricing", platform.openai.com/docs/pricing (the GPT-5.6 Sol standard, batch and flex rates, the cached input rate, the sentence on promotional pricing "at least through November 21, 2026", and the sentence on aliases); Meta, "Pricing", WhatsApp Business Platform developer documentation (South Africa listed under Rest of Africa, calling code 27; rates served interactively), re-read this morning against our 2 September reading; the 1 October 2026 service message change as quoted from Meta's documentation in our 2 September post. Every quoted phrase is the publisher's own wording. Every figure is in US dollars per million tokens unless stated, and none of them is a NexBDM price.
Published on nexbdm.agency. Want this applied to your business? Run the free Autopsy diagnostic →