Founders do not think in tokens. They think in deliverables, and they approve budgets that express cost per blog post, per ad variant, per report: against what those deliverables cost today. Reframe your AI budget in those units, show the before-and-after, and add a hard guardrail against runaway spend, and the conversation stops being a negotiation and becomes an obvious yes.
Key Takeaways
- Budget in cost per deliverable, never cost per million tokens. That is the unit your founder already uses.
- Benchmark your current cost per deliverable first; without it you have no case, only a request.
- A small frontier budget plus a large cheap-model budget usually beats a mid-tier compromise at the same total.
- API access is typically far cheaper than SaaS seats that mark up the same underlying models: but not always, once implementation time is counted.
- One page, three numbers: current cost, proposed cost, and the cap that prevents a surprise invoice.
- All modelled figures use list prices as of September 2026, check current pricing pages.
If it does not fit on one page, it will not get read before the meeting.
Why Token Budgets Get Rejected
A request for "$400 a month in API credits" invites exactly one question: for what? And the honest answer, "tokens", is not an answer a founder can evaluate. There is no benchmark, no comparison, no unit they recognise.
A request for "roughly $23 per published blog post, against the $120 we currently spend, for 20 posts a month" is self-evaluating. The founder does the arithmetic in their head and either agrees or asks about quality, which is a productive conversation.
Step One: Benchmark What You Spend Today
You cannot propose a new cost per deliverable without knowing the current one.
Count the full cost, honestly
For each deliverable type, include:
- Freelance or agency invoices
- Internal time at a loaded hourly rate (salary plus overhead, not salary alone)
- Tool subscriptions attributable to that output
- Review and approval time from anyone who touches it
Use published units, not drafted ones
Divide by what actually shipped. Abandoned work is a real cost and it belongs on the pieces that survived.
Expect the number to be uncomfortable
Most teams discover their true cost per blog post is two to four times what they assumed, because internal time was never counted. That discomfort is your strongest argument.
Step Two: The Cost-Per-Deliverable Table
This is the centre of the business case. All figures modelled from list prices as of September 2026, check current pricing pages. Human time assumed at a loaded $30/hour; substitute your own rate and rerun.
| Deliverable | Current all-in cost | Proposed model tier | Token cost | Human time | Proposed all-in | Monthly volume | Monthly line |
|---|---|---|---|---|---|---|---|
| 1,500-word blog post | $120 | Frontier + caching | ~$0.45 | 45 min | ~$23 | 20 | $460 |
| Ad variant (single creative copy set) | $18 | Flash-class | ~$0.01 | 6 min | ~$3 | 120 | $360 |
| Weekly performance report | $90 | Frontier | ~$0.30 | 25 min | ~$13 | 4 | $52 |
| Social repurpose (1 post → 6 assets) | $45 | Flash-class | ~$0.02 | 15 min | ~$7.50 | 30 | $225 |
| Landing page copy | $260 | Frontier | ~$1.10 | 90 min | ~$46 | 3 | $138 |
| Email sequence (5 emails) | $220 | Frontier + caching | ~$0.80 | 70 min | ~$36 | 2 | $72 |
Totals: current-equivalent monthly cost roughly $8,235; proposed roughly $1,307. API spend within that: under $40.
How to read this table with your founder
Point at the token cost column. It is under $40 a month across every deliverable. Then point at the human time column, which is over $1,250. The conversation you are having is about production capacity, not about software spend, and that framing is what gets it approved.
Be upfront about what is modelled
These are estimates built from list prices and plausible editing times, not measured results from your team. Say so. A business case that admits its assumptions survives scrutiny; one that presents models as guarantees does not survive the first month of actuals.
Per deliverable is the only unit that travels outside the marketing team.
Step Three: The Build-vs-Buy Question
Your founder will ask why you cannot just use the tool you already pay for. Here is the honest comparison.
What SaaS seats actually charge for
Most AI marketing tools are a wrapper around the same frontier models: GPT-6 Astra or Claude Fable 5.1 at $10/M input and $50/M output, list prices as of September 2026. The seat price covers the model, a UI, templates, and a margin.
Work the method on a typical seat:
Assumption: $60/seat/month, 5 seats = $300/month. Underlying token consumption for the same work: roughly $40/month. Implied markup: 7.5x.
When the markup is worth paying
- Non-technical team with no one to build workflows
- You need the governance, audit trail or approvals the tool provides
- The tool does something genuinely beyond generation: scheduling, publishing, analytics integration
- Volume is low enough that engineering time exceeds the markup
When it is not
- You have someone who can wire up API calls
- Your volume is meaningful
- You need custom brand context in every generation, which most tools handle poorly
- You want caching, which seat-based tools rarely expose
The realistic answer for most teams
Both. Seats for the people who need a UI, API access for the workflows that run at volume. Presenting it as either/or invites a false choice.
Step Four: The Barbell Allocation
This is the recommendation most budgets get wrong.
The instinct
Pick one mid-tier model, standardise everyone on it, budget a single number. Simple, defensible, and usually the worst value available.
Why it fails
Mid-tier is too expensive for the 80% of work that is low-stakes volume, and not good enough for the 20% that carries real consequence. You overpay for the bulk and underperform on the work that matters.
The barbell
- Small frontier budget: GPT-6 Astra or Claude Fable 5.1, with caching where context repeats, for public-facing long-form, strategy, and anything a prospect will judge you by
- Large cheap-model budget: Gemini 3.8 Flash for outlines, metadata, bulk repurposing, internal summaries
Flash-class pricing has been collapsing: Gemini 3.7 Flash launched at half the per-million-token price of 3.6 Flash from three weeks earlier, and 3.8 Flash arrived in early September as Google's best available model. That makes the cheap end of the barbell better value every quarter. Check ai.google.dev/pricing for current rates.
A note on caching in the budget
Anthropic's Claude Fable 5.1 cut cache reads by roughly 75% to about $0.25/M, and claims around 25% cost reduction on typical workloads with up to 45% on agentic tasks. Since GPT-6 Astra and Claude Fable 5.1 carry identical headline pricing, caching behaviour is a legitimate line in your build-vs-buy reasoning, not a detail. Verify at anthropic.com/pricing and openai.com/pricing.
Step Five: The One-Page Business Case
The structure that works
- Current state: one table, cost per deliverable today, total monthly
- Proposed state: same table, proposed costs, total monthly
- The delta: one number, in currency, per month
- What we are not claiming: quality is maintained by keeping human editing in the loop; this is a cost and capacity change, not a headcount replacement
- The guardrail: the hard cap, stated as a number
- Review date: 90 days, with the actuals-versus-model comparison committed to up front
What to leave out
Model names in the headline. Token arithmetic in the body. Anything that requires your founder to learn a new unit to evaluate the proposal. Put the technical detail in an appendix for the one person who will ask.
The sentence that closes it
"We will review actuals against this model in 90 days, and if the per-deliverable cost is more than 25% above what I have projected, we cut the scope rather than asking for more."
One page, one delta, one cap. The rest is appendix.
Step Six: Guardrails Against a Surprise Invoice
The fastest way to lose an approved AI budget is one runaway agent producing a five-figure bill. Set these before you need them.
Hard spend caps at the provider
Every major provider supports monthly limits. Set one. Set it slightly above your model, not at your model, so a busy week does not halt production.
Per-run token ceilings
Any agentic workflow gets a maximum token budget per execution. If it exceeds it, it stops and alerts rather than continuing.
Maximum call counts per workflow
An agent stuck in a critique-revise loop can make hundreds of calls. Cap the loop count explicitly.
Expiry on every scheduled job
Scheduled experiments must have an end date. An experiment without one is a permanent subscription nobody approved.
Weekly attribution, not monthly
Tag calls by workflow and review weekly. A runaway caught on day three costs a fraction of one caught on day twenty-eight.
Alert on rate of change, not on total
An alert at 80% of budget fires too late. An alert on daily spend exceeding 3x the trailing average catches problems while they are small.
Common Objections and Honest Answers
"Can't we just do this with free tools?" For low volume, largely yes. The case for paid access is volume, consistency and the ability to inject full brand context, which free interfaces make tedious.
"What if prices go up?" They have been falling, sharply and repeatedly, through 2026. But the budget should survive a rate increase, which is why the guardrail is a spend cap rather than a token quota.
"Does this replace a hire?" No, and do not claim it does. It changes what one person can produce. Presenting it as headcount replacement is the fastest way to lose credibility when output quality needs human judgement, which it does.
"Why not the newest model everyone is talking about?" Some of what is being talked about is not available to you. Claude Mythos is gated to vetted cybersecurity and life-sciences professionals. Gemini 3.5 Pro has not shipped. And OpenAI's Astra is unrelated to Google's Project Astra, a research prototype that never became a product.
For broader benchmarks on content operations and marketing spend, HubSpot's marketing statistics library is a reasonable external reference: hubspot.com/marketing-statistics.
Frequently Asked Questions
How much should a marketing team budget for AI in 2026? Derive it from deliverables rather than picking a number. For a team publishing 20 posts and 120 ad variants monthly, direct API spend is often under $50; the bigger line is the human time that surrounds it.
Why budget per deliverable instead of per token? Because founders, finance teams and everyone outside marketing already think in deliverables. Per-token budgets cannot be evaluated by anyone who does not work with models daily.
Is API access cheaper than SaaS seats? Usually on raw cost, markups of 5–10x over underlying token cost are common, but seats can be cheaper once implementation and maintenance time is counted at low volume.
What is the barbell approach to AI budgeting? A small frontier-model budget for high-consequence work plus a large cheap-model budget for volume, instead of standardising everyone on one mid-tier model.
How do I stop a runaway agent from blowing the budget? Provider-level spend caps, per-run token ceilings, maximum call counts per workflow, expiry dates on scheduled jobs, and alerts on rate of change rather than on total.
What should the one-page business case contain? Current cost per deliverable, proposed cost per deliverable, the monthly delta, an explicit statement of what you are not claiming, the spend cap, and a 90-day review date.
Should I include human editing time in the budget? Yes, always. It is the largest component, and omitting it makes the proposal look dishonest the moment anyone checks.
How often should I revisit the budget? Every quarter. Prices moved dramatically inside three-week windows during 2026, a model allocation set a year ago is almost certainly no longer optimal.
What if my actuals come in above the model? Say so early and cut scope rather than requesting more. Credibility on the first budget determines how easily you get the second.
Do I need to name specific models in the business case? Not in the body. Put them in an appendix. The founder is approving a cost per deliverable, not a vendor.
Let's Talk
If you are putting this case together and want someone to pressure-test the numbers before it reaches your founder, I am happy to help. My work and the contact form are both at younusfardeen.com: 4+ years of marketing experience helping edtech and startup brands grow organically, usually on budgets that had to be justified line by line first.