Prompt caching lets a model charge you far less to re-read context it has already seen recently, which means the brand guidelines, tone rules and product catalogue you paste into every generation can become almost free instead of being billed at full input price every single time. The practical move is simple: put your stable brand context at the very start of the prompt and the variable brief at the end. As of September 2026, Anthropic's Claude Fable 5.1 cut cache reads by roughly 75% to $0.25 per million tokens, which changes the economics of context-heavy marketing workflows overnight.
Key Takeaways
- Caching bills repeated context at a steep discount instead of full input price, Claude Fable 5.1 cache reads sit at $0.25/M versus $10/M standard input (list prices as of September 2026, check current pricing pages).
- Prompt order is everything. Caching works on a stable prefix, so anything that changes must sit after everything that doesn't.
- This makes "feed the model our entire brand bible every time" economically sane for the first time.
- Caching does nothing for genuine one-offs or workflows where context changes on every call.
- The quality upside is bigger than the cost upside: teams stop trimming context to save money and start giving the model everything it needs.
- Restructuring your prompt library is a half-day job with a payback measured in weeks, not quarters.
A cached prompt behaves like a foundation: build it once, reuse it on every generation.
What Prompt Caching Actually Is
When you send a prompt to a language model, the model processes every token from scratch. That processing is what you pay for on the input side. Prompt caching changes the deal: if the opening chunk of your prompt is byte-for-byte identical to one the provider processed for you recently, it can reuse the processed state instead of redoing the work, and it charges you a fraction of the normal rate for that portion.
Think of it as the difference between re-reading a 40-page brand guide before every task versus having already read it and only needing the new brief.
The two prices you need to know
Caching introduces two extra line items beyond standard input and output:
- Cache write, a small premium the first time a block is written into cache.
- Cache read, a heavily discounted rate every subsequent time that block is reused.
The write premium is why caching a prompt you use twice is roughly break-even, and why caching a prompt you use two hundred times is transformative.
Why this landed now
Both major frontier models sit at the same headline price as of September 2026: GPT-6 Astra (released 3 September 2026) and Claude Fable 5.1 (released 1 September 2026) are each $10 per million input tokens and $50 per million output tokens. When the sticker price is identical, the differentiator moves to the fine print, and Anthropic's cache read drop to roughly $0.25/M is the loudest piece of fine print in the category. Anthropic claims around 25% cost reduction on typical workloads and up to 45% on agentic tasks. Verify against anthropic.com/pricing and openai.com/pricing before you budget against it.
What Qualifies for Caching
Not everything you send is cacheable. In practice, three conditions have to hold.
It has to be a prefix
Caching matches from the start of the prompt forward. The moment your text diverges from the cached version, the cache stops applying. One changed word near the top invalidates everything after it.
It has to be big enough
Providers set minimum cacheable block sizes. A two-line tone reminder will not qualify. A full style guide, a product catalogue, or a set of ten past-campaign examples comfortably will.
It has to be reused within the window
Caches expire. If your team generates in bursts, a content sprint on Tuesday morning, you get excellent hit rates. If one person runs one prompt a day, you will mostly pay cache-write prices with no reads to amortise them.
Why Prompt Order Matters More Than Prompt Wording
Most marketing prompt libraries were written by people optimising for readability, which usually means brief first, context second. That is exactly backwards for caching.
The default (uncacheable) structure
Write an Instagram carousel for our new data analytics course
launching 22 September, aimed at working professionals in
tier-2 cities.
Here is our brand voice guide: [3,000 words]
Here are our product details: [2,000 words]
Here are five past posts that performed well: [1,500 words]Every generation changes the first paragraph, so nothing caches. You pay full input price on 6,500 words of static context, forever.
The cached structure
[SYSTEM / STABLE BLOCK, identical on every call]
Brand voice guide: [3,000 words]
ICP definitions: [800 words]
Product catalogue: [2,000 words]
High-performing past campaigns: [1,500 words]
Formatting and compliance rules: [400 words]
[VARIABLE BLOCK, changes per task]
Task: Write an Instagram carousel for our new data analytics
course launching 22 September, aimed at working professionals
in tier-2 cities.Identical information. Radically different bill. The stable block is now a cache prefix, and the only fully-priced input is the two-sentence brief.
The before/after is not about writing better prompts. It is about ordering the same prompt correctly.
Doing the Arithmetic Honestly
Here is the method, so you can rerun it with your own numbers rather than trusting mine.
Assumptions: 6,500 words of stable context ≈ 8,700 tokens. 60 generations per week. List prices as of September 2026, check current pricing pages.
Without caching: 8,700 tokens × 60 runs = 522,000 input tokens/week 522,000 ÷ 1,000,000 × $10 = $5.22/week on context alone
With caching (one write, 59 reads): Write: 8,700 ÷ 1,000,000 × $10 (plus provider write premium) ≈ $0.09–0.12 Reads: 8,700 × 59 = 513,300 ÷ 1,000,000 × $0.25 = $0.13 Total ≈ $0.25/week
That is a modelled figure, not a guarantee. Cache hit rates in the real world are never 100%, write premiums vary by provider, and your context is probably a different size than my example. Rerun it with your own token counts.
The number that matters is not the dollar saving
At this volume you saved five dollars a week. That is noise. The real change is that the cost objection to sending full context disappears, and full context is what makes output stop reading like generic AI copy.
How This Changes the Brand Bible Question
For three years the standard advice has been "keep prompts tight." That advice was cost-driven, not quality-driven. Teams trimmed their brand guidelines to a paragraph because a full guide was expensive to send 200 times a month.
What you can now afford to include
- The complete tone-of-voice document, not a summary of it
- Banned phrases and competitor-mention rules in full
- Ten to twenty annotated examples of work that hit the mark
- Full ICP definitions with objections, not one-line personas
- Actual product specs, pricing tiers and eligibility rules
What changes in the output
In edtech work specifically, where claims about outcomes, placements and eligibility carry real consequences, the difference between a summarised context block and a complete one is the difference between output you rewrite and output you edit. Over a year of running content for Masai School, the single biggest driver of usable first drafts was never the model choice. It was how much real, specific context the model had to work from.
When Caching Doesn't Help
Being honest about the limits is what separates this from vendor marketing.
Genuine one-off tasks
Ad-hoc analysis, a single competitor teardown, a one-time positioning exercise. You pay the cache write premium and never read it back. Slightly worse than not caching.
Constantly-changing context
If the bulk of your prompt is a fresh document every time, analysing this week's search console export, summarising today's customer calls, there is no stable prefix to cache. The variable part is the whole prompt.
Low-frequency, scattered usage
One marketer running four prompts a day across four different projects will see poor hit rates. Caching rewards concentration and cadence.
Workflows where the context genuinely should change
Do not freeze a context block just to preserve a cache hit. If your positioning shifted, update the block. The saving is not worth shipping stale messaging.
Setting Your Team Up in an Afternoon
Step one: audit what repeats
Pull twenty recent prompts from your team. Highlight every sentence that appeared in more than five of them. That highlighted text is your cache block.
Step two: build one canonical stable block
Consolidate into a single versioned document. Version it, brand-context-v4, so you know when a change invalidated the cache.
Step three: rewrite your templates
Every template becomes: stable block, then a separator, then the brief. Nothing else.
Step four: measure hit rate, not spend
Most providers report cache hit statistics. A hit rate below 60% usually means someone is editing the stable block, or your cache window is expiring between sessions.
One canonical context block, versioned and shared, beats twelve personal prompt variations.
Where the Other Models Sit
Caching is not exclusive to one vendor, but the aggressiveness of the discount varies. GPT-6 Astra prices cache reads and writes separately: check openai.com/pricing for the current figures, and note that its "fast mode" runs at twice the speed for twice the price, which compounds any inefficiency in your prompt structure.
The Flash-class tier is a different calculation entirely. Gemini 3.7 Flash launched at half the per-million-token price of Gemini 3.6 Flash just three weeks earlier, and Gemini 3.8 Flash arrived in early September as Google's best available model. When the base rate collapses that fast, caching matters less as a cost lever and more as a latency one. Current rates are at ai.google.dev/pricing.
Open-weight options, Qwen3.8 27B (2 September) and Z.AI GLM-5.3 Flash (26 August), shift the question from per-token price to infrastructure cost, where caching becomes a memory-management decision rather than a billing one.
One clarification worth making because the naming confuses people: OpenAI's Astra has nothing to do with Google's Project Astra, which was a research prototype and never shipped as a product. And Claude Mythos is gated to vetted cybersecurity and life-sciences professionals. It is not an option for marketing teams, whatever you read in a roundup post.
A Note on Context Windows
You will see confident numbers quoted for the context windows of GPT-6 Astra and Claude Fable 5.1. As of this writing, neither has been published. Anyone giving you a figure is guessing. Plan your context strategy around what you actually need to include, not around a ceiling nobody has confirmed.
Frequently Asked Questions
What is prompt caching in simple terms? It is a discount on repeated context. If the opening section of your prompt matches one the model processed recently, you are charged a fraction of the normal input rate to reuse it.
How much can prompt caching actually save a marketing team? It depends entirely on how much static context you send and how often. Teams sending a large brand context block dozens of times a week can see the context portion of their input bill drop by 90% or more. Teams doing scattered one-offs will see almost nothing.
Does prompt caching change the quality of the output? Not directly: the model sees the same tokens either way. Indirectly it matters a lot, because it removes the cost incentive to send less context than the task needs.
Why does prompt order affect caching? Caches match from the beginning of the prompt forward. Once your text differs from the cached version, everything after that point is processed fresh. Put stable content first so the divergence happens as late as possible.
How long does a cache last? Cache lifetimes vary by provider and are short: think minutes, not days. Check the current documentation on the provider's pricing page rather than relying on a figure in a blog post.
Should I cache my system prompt or my user message? Whichever holds the stable content. In most marketing setups, the system prompt is the natural home for brand context, but the mechanism cares about position, not role.
Is caching worth it if I use a cheap Flash-class model? Less so. When the base input rate is already very low, the absolute saving shrinks. Caching on Flash-class models is usually about response latency rather than cost.
Does editing one word in my brand guide break the cache? Yes, for everything after that word. That is why you version the block and make changes deliberately rather than continuously.
Can I cache images or documents, not just text? Support varies by provider and model. Check the current documentation. This is an area changing month to month.
What is the single fastest change to make today? Move your brand context above your brief in every template. That one reordering captures most of the benefit before you optimise anything else.
Work With Me
If you want a second pair of eyes on how your team structures prompts, or on the wider organic growth engine they feed, you can see my work and get in touch through the contact form at younusfardeen.com. I have spent 4+ years in marketing helping edtech and startup brands grow organically, and I am always happy to talk through what is working and what isn't.