Per-token prices fell across every tier through 2026: Gemini 3.7 Flash launched at half the price of its predecessor three weeks earlier, and Anthropic cut Claude Fable 5.1 cache reads by roughly 75%, yet a lot of marketing teams are looking at larger invoices than last quarter. The reason is consumption, not price: agentic workflows make dozens of model calls where chat made one, reasoning models bill you for thinking you never read, and bigger context windows invite stuffing more in.
Key Takeaways
- Unit prices falling while total spend rises is a volume story, not a pricing story.
- A single agentic deliverable can consume 20–50x the tokens of an equivalent chat exchange.
- Reasoning tokens are billed output you never see, and they scale with problem difficulty.
- Larger context windows create a "might as well include it" habit that quietly multiplies input.
- Diagnosis requires per-workflow attribution, not a single account-level total.
- Routing discipline, not model-switching, is what actually brings the bill down.
Falling rate, rising bill. Both charts are true at once.
The Arithmetic That Explains Everything
Spend = volume x rate.
Rates dropped maybe 30–50% year on year at most tiers. Volume, for teams that adopted agentic tooling, went up by considerably more than that. Multiply a 0.6 rate multiplier by a 15x volume multiplier and you get a 9x bill.
Nothing about this is a billing error or a vendor trick. It is what happens when a capability becomes affordable enough to use constantly.
Cause One: Agentic Workflows Eat Tokens
What changed in how we use models
A chat exchange is one call. You write a brief, the model writes a draft, you are done. An agentic workflow for the same deliverable might:
- Plan the approach
- Search or retrieve source material
- Read and summarise each source
- Draft an outline
- Critique the outline
- Revise the outline
- Draft the full piece
- Critique the draft against the brief
- Revise
- Fact-check specific claims
- Format and generate metadata
Eleven calls minimum, and several of them resend the growing document as input. The output of step seven becomes the input of steps eight, nine and ten.
The compounding effect
Input grows monotonically through an agent run. By call ten you may be sending 30,000 tokens of accumulated context to produce 500 tokens of output. This is why Anthropic's claim of up to 45% cost reduction on agentic tasks from cheaper cache reads is specifically about this pattern, agentic chains are where repeated context is most severe.
Why teams do not notice
Because the workflow feels like one action. You click run, you get a deliverable. The dashboard shows one task. The invoice shows forty calls.
Cause Two: You Are Paying for Reasoning You Cannot Read
Reasoning models generate extended internal deliberation before producing a final answer. That deliberation is output tokens. It is billed at the output rate, which at frontier level is $50 per million as of September 2026 for both GPT-6 Astra and Claude Fable 5.1.
The scaling problem
Reasoning length scales with perceived problem difficulty. A vague brief invites more deliberation than a precise one. A poorly-specified task can cost several times a well-specified one for an identical final output length.
The practical implication
Brief quality is now a cost lever, not just a quality lever. "Write something good about our new course" and "Write a 1,200-word launch announcement for the data analytics course, following the attached structure, for working professionals in tier-2 cities" produce different bills.
You are billed for the crumpled pages too.
Cause Three: Context Stuffing
Larger context windows removed the discipline that scarcity used to impose.
When context was tight, you chose what mattered. When it is generous, the path of least resistance is to include everything, the full transcript, the whole competitor site, every past campaign, because filtering takes effort and inclusion does not.
Why this is worse than it looks
Padding context with marginally relevant material does not just cost input tokens. It degrades output, because signal-to-noise falls and the model weights irrelevant material. You pay more and get less.
A discipline worth restoring
For each context element, ask: would removing this change the output? If you cannot answer confidently, run it both ways once. Most teams find a third of their context is inert.
One caution: the context window sizes for GPT-6 Astra and Claude Fable 5.1 have not been published. Do not design a context strategy around a ceiling figure you read somewhere, design it around what the task genuinely needs.
Cause Four: "Just Try It On Everything"
Every new model release triggers a wave of experimentation. Someone runs the whole content backlog through it to compare. Someone else builds an agent to test whether it can handle reporting. A third person leaves a scheduled job running after the experiment ended.
This is healthy behaviour with an unhealthy billing profile, and it is almost entirely unmeasured because experiments rarely have owners or budgets.
The specific failure mode
An experiment that runs on a schedule and is never switched off. It produces output nobody reads and a line item nobody questions, because the total looks plausible.
Diagnosing Where Your Spend Actually Goes
You cannot fix what you cannot attribute. Here is the sequence.
Step one: tag every call
Most providers support metadata or separate API keys per project. Use them. One key per workflow: content drafting, social repurposing, reporting, research, experiments. Without this, all diagnosis is guesswork.
Step two: compute cost per deliverable
For each workflow, divide monthly spend by units produced. Blog posts, ad variants, reports, whatever the workflow outputs. This immediately surfaces the workflows that are expensive per unit rather than merely large in total.
Step three: split input, output and cache
Three very different problems:
- Input-heavy means context stuffing or missing caching
- Output-heavy means verbosity or excessive reasoning
- Low cache hit rate means your prompt structure is wrong
Step four: find the calls with no consumer
Trace each workflow to a human or a system that uses its output. Anything that terminates in nothing is pure waste. In most audits I have run, this is 10–20% of spend.
Step five: rank by spend, fix top three
Do not optimise everything. The distribution is almost always heavily skewed, two or three workflows account for most of the bill.
Attribution first. Optimisation second. In that order, or you will optimise the wrong thing.
Setting Per-Workflow Budgets
Budget by deliverable, not by month
"₹40 per published blog post" is enforceable and diagnosable. "₹30,000 a month for AI" is neither, it tells you nothing when you exceed it.
Set the number from a measurement, not a hope
Measure current cost per deliverable for four weeks. Set the budget 20% below it. That gap is the optimisation target.
Make the budget visible where the work happens
A budget in a finance spreadsheet changes nothing. A budget shown in the tool, or reported weekly to the team running the workflow, changes behaviour.
Handle experiments separately
Give experimentation a fixed monthly allowance with a hard expiry on every job. An experiment without an end date is a subscription.
The Routing Discipline That Fixes It
This is the single highest-leverage change, and it costs nothing to implement.
The principle
Match model tier to consequence of error, not to habit.
A workable routing table
- Frontier (GPT-6 Astra, Claude Fable 5.1): public-facing long-form, strategy documents, anything where a subtle error damages credibility
- Frontier with caching, any high-volume task carrying a large stable brand context; the cache read rate makes this cheaper than it looks
- Flash-class (Gemini 3.8 Flash): outlines, meta descriptions, alt text, bulk social repurposing, internal summaries
- Open-weight self-hosted (Qwen3.8 27B, GLM-5.3 Flash), only at high sustained volume with someone to own the infrastructure
- Nothing at all: a genuine option, and an underused one
The tier that usually wins
Teams default to a mid-tier compromise model for everything. This is almost always the worst allocation. A small frontier budget for work that matters plus a large cheap-model budget for volume beats a single mid-tier spend at the same total.
Two things to stop doing
Stop running fast mode by default. GPT-6 Astra's fast mode is twice the speed at twice the price; for batch content work you are paying double for a latency improvement nobody experiences.
Stop chasing models you cannot use. Claude Mythos is gated to vetted cybersecurity and life-sciences professionals. It is not a marketing option regardless of what a comparison post implies. And OpenAI's Astra is unrelated to Google's Project Astra, a research prototype that never shipped; conflating them leads to genuinely wrong routing decisions.
What Good Looks Like After Ninety Days
In the content operations I have worked on: including three years running organic growth for Masai School, where Instagram went from 26K to 117K and LinkedIn from 50K to 160K, the pattern that holds is unglamorous: attribution, then routing, then caching, then brief quality. Model switching comes last and matters least.
A realistic target is 40–60% reduction in spend with no reduction in output volume, achieved almost entirely by sending the right work to the right tier and switching off jobs nobody reads.
For current rates as you rebuild your routing, go to primary sources: openai.com/pricing, anthropic.com/pricing, and ai.google.dev/pricing. List prices as of September 2026 change frequently, check them rather than trusting any figure quoted here.
Frequently Asked Questions
Why is my AI bill going up when prices are falling? Because your consumption is rising faster than rates are falling. Agentic workflows, reasoning tokens and context stuffing all increase volume substantially.
How many more tokens does an agentic workflow use than chat? It varies by design, but 20–50x for a comparable deliverable is a realistic range, driven mostly by repeatedly resending accumulated context.
Are reasoning tokens billed even though I cannot see them? Yes. Internal reasoning is output, billed at the output rate. Vaguer briefs tend to generate longer reasoning.
What is the fastest way to cut my AI spend? Attribute spend per workflow, then find the calls whose output nobody consumes. That alone is typically 10–20% of the bill.
Does prompt caching actually help with agentic costs? Yes: agentic chains resend large stable context repeatedly, which is exactly the pattern caching targets. Anthropic claims up to 45% cost reduction on agentic tasks with Claude Fable 5.1's cheaper cache reads.
Should I just switch to a cheaper model? Not wholesale. Route by consequence of error. Cheap models on high-stakes public content cost more once editing time is counted.
How do I stop a runaway agent from producing a surprise invoice? Hard per-run token caps, a maximum call count per workflow, and expiry dates on every scheduled job. Set these before you need them.
Is a bigger context window a good thing? It is a capability, not an instruction. Include what changes the output; exclude the rest. More context often degrades quality as well as costing more.
How should I budget for AI across a marketing team? Per deliverable, not per month. Cost per blog post, per ad variant, per report, numbers a team can act on and a founder can understand.
How often do these prices change? Frequently enough that any figure more than a month old should be re-verified. A Flash-class price halved within three weeks during 2026.
Work With Me
If your AI spend has quietly outgrown the results it produces, that is usually a routing and attribution problem rather than a vendor problem: and it is fixable in a few weeks. You can look through my work and reach me via the contact form at younusfardeen.com. Four-plus years in marketing, most of it spent helping edtech and startup brands grow organically without overspending to do it.