To evaluate AI models for marketing properly, you need a fixed set of ten real briefs from your own workflow, blind scoring by people who didn't run the test, and a stopwatch measuring how long an editor takes to make each output publishable. That's the whole protocol. Public benchmark scores, FrontierMath, ARC-AGI, whatever ships next quarter, tell you close to nothing about whether a model will write your landing page copy well, because they measure capabilities that don't overlap with marketing work and they're reported by the vendors selling the model.
This is a two-day exercise. I've run it with several teams and it has never once produced the answer people expected going in.
Key Takeaways
- Build a fixed eval set of 10 real tasks from your actual workflow. Reuse it every time you evaluate a model.
- Score blind, strip model labels before anyone reads the output.
- The five dimensions that matter: brief adherence, voice match, factual reliability, format compliance, and editing time.
- Editing time is the real cost metric. An editor's hour dwarfs a million output tokens.
- Vendor benchmark claims are self-reported and measure the wrong things for marketing.
- Re-run the eval quarterly. Model releases in 2026 are fast enough that annual decisions go stale.
Why Public Benchmarks Don't Help You
As of September 2026 the flagship models are OpenAI's GPT-6 Astra (released 3 September) and Anthropic's Claude Fable 5.1 (released 1 September). Both are priced at $10/M input and $50/M output. OpenAI has reported figures for Astra including 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 47% faster computer use.
Those are OpenAI's own numbers. Set aside whether they're accurate: assume they are. Now ask: what does a FrontierMath score predict about how well a model writes a product-led blog post that hits a keyword, respects a 40-rule style guide, and doesn't invent a statistic?
Nothing. There's no established correlation, and there's no reason to expect one. Competitive mathematics and brand-voice copywriting are not adjacent skills.
Three specific problems with benchmarks
They're self-reported. Vendors run them, vendors publish them, vendors choose which to publish. Attribute these numbers to the vendor. Never state them as fact.
They're saturated at the top. When a model scores 99.9%, the benchmark has stopped discriminating. It cannot tell you which of two models is better because both max it out.
They test capability ceilings, not typical performance. Marketing work almost never sits near a frontier model's ceiling. What matters is consistency at the median task, which benchmarks don't measure at all.
What benchmarks are actually useful for
Rough tier placement. If a model scores catastrophically badly across the board, skip it. Beyond that filter, they're noise.
Building Your Eval Set
Pick ten real tasks
Not ten tasks you invent for the test. Ten briefs you actually ran in the last quarter, with the actual inputs and the actual published output as a reference point.
A reasonable spread:
- A long-form blog brief with keyword targets and a style guide
- A short product description from structured product data
- Five ad copy variants from one value proposition
- A LinkedIn post in a founder's voice
- An email subject line set (10 variants)
- A summarisation task, long document to 200 words
- An extraction task, pull structured fields from unstructured text
- A rewrite, take an off-voice draft and bring it on-voice
- A meta description and title tag set for five URLs
- A complex multi-constraint brief, the hardest thing you routinely ask for
That spread matters. Models differ enormously by task type, and an eval set that's all long-form will tell you nothing about whether you can route bulk work to a cheap tier.
Freeze it
Write the briefs down verbatim. Save the inputs. Version the file. The value of the eval set is that it's the same every time, that's what lets you compare a September result against a December result.
Include a human reference
For each task, include the human-written version you actually published. This gives your scorers an anchor and lets you ask the most useful question in the whole protocol: is the model output above or below the bar we already ship at?
The Scoring Rubric
Use this table. Score every output on every dimension. Weight according to what your team actually cares about, but start here.
| Dimension | What you're measuring | Scale | Weight | How to score it |
|---|---|---|---|---|
| Brief adherence | Did it do what the brief asked, all of it? | 1-5 | 30% | Count constraints in the brief; score by proportion met. Objective. |
| Voice match | Does it sound like your brand? | 1-5 | 20% | Compare against the human reference. 3 = indistinguishable. |
| Factual reliability | Any invented facts, stats, names, or claims? | 1-5 | 20% | 5 = nothing invented. Any fabricated statistic caps the score at 2. |
| Format compliance | Correct structure, headings, length, markdown? | 1-5 | 10% | Binary per requirement; average them. |
| Editing time | Minutes to make it publishable | Minutes | 20% | Stopwatch. Single editor. No interruptions. |
| Cost per task | Token cost of the run | Currency | Tiebreak | From the API response. Record but don't over-weight. |
Notes on each dimension
Brief adherence is the highest-weighted item because it's the thing that actually differentiates models in practice and it's the most objectively scorable. Count the constraints. A brief that says "800-1000 words, include these three keywords, cite two sources, end with a CTA, no first person" has six constraints. Score the proportion met.
Voice match is subjective, which is why it needs the human reference and blind conditions.
Factual reliability deserves the harsh cap. A model that writes beautifully and invents a statistic has produced negative value, it costs more to fact-check than to write from scratch.
Editing time is the one people skip and the one that decides the outcome. Time it honestly. One editor, one timer, no phone.
Cost per task is a tiebreak, not a headline. If two models score within noise of each other, take the cheaper one. If one scores meaningfully higher, cost rarely justifies choosing the worse output: unless the volume is enormous, in which case see the routing section.
Running the Test Blind
Why blind matters
If your scorer knows which output came from the expensive flagship model, they will score it higher. This is not a character flaw; it's how everyone works. Blind conditions are the only way to get a signal.
The mechanics
- Run all 10 briefs through every model under test. Same prompts, same parameters, same day.
- Save outputs to files named by random ID, with a separate key file only you hold.
- Strip any model-identifying artefacts: some models have telltale formatting habits, so normalise where you can.
- Give the randomised set to two or three scorers who know your brand and did not build the prompts.
- Collect scores. Then reveal the key.
Who should score
People who edit your content day to day. Not the person who ran the test, not the person championing a particular vendor, not the CMO. The people who'll live with the output.
Measuring Editing Time Properly
This deserves its own section because it's where the money is.
The method
One editor. One output at a time. A timer started when they open the file and stopped when they'd be willing to publish it. No multitasking. Record the number even when it's embarrassing.
Turning it into money
Take your editor's loaded hourly cost. Multiply by the time difference between models across a realistic monthly volume. Compare against the token cost difference.
A worked example: two models differ by 8 minutes of editing per piece. You publish 100 pieces a month. That's 13.3 hours. At a loaded rate of, say, $40/hour, that's $533/month. Token cost differences at typical marketing volumes rarely approach that. Which is why the cheap-model argument only holds where editing time is genuinely equivalent: usually on short, structural, high-volume work.
The inverse case
Sometimes the cheap model's editing time is identical: on meta descriptions, extraction, bulk variants. There, every token you spend on a flagship model is pure waste. The eval set is what tells you which category each task falls into.
What to Test in September 2026
The current field, with the caveats that matter:
- GPT-6 Astra (OpenAI, 3 September 2026). $10/$50, fast mode at 2x speed and 2x price. Off by default for enterprise admins at rollout: check you can actually access it before planning a test. Available on ChatGPT Plus/Pro/Business/Enterprise, the API, Azure and AWS Bedrock, with a Pro variant on higher tiers.
- Claude Fable 5.1 (Anthropic, 1 September 2026). Same $10/$50 headline. Cache reads cut roughly 75% to $0.25/M; Anthropic claims typical savings around 25%, up to 45% agentic. Lineup runs Mythos > Fable > Opus > Sonnet > Haiku. Mythos is gated to vetted cybersecurity and life-sciences professionals: marketers can't test it, so don't put it in your comparison.
- Gemini 3.8 Flash (Google). The best available Gemini model as of early September 2026. Gemini 3.5 Pro has not shipped. Four Flash models in 106 days, with 3.7 Flash launching at half the per-token price of 3.6 Flash three weeks earlier.
- Open-weight: Qwen3.8 27B (2 September 2026, self-hostable), Z.AI GLM-5.3 Flash, DeepSeek V4 Flash.
Neither OpenAI nor Anthropic has published context window figures for their September flagships. If context length matters to your workflow, test it rather than citing a number.
All pricing here is as of September 2026 and moves. Check OpenAI, Anthropic and the Google blog before you budget anything.
Turning Results Into a Routing Decision
The output of a good eval isn't "we use model X." It's a routing table: task type mapped to model tier, based on where the quality difference justifies the price difference.
Most teams find their eval splits cleanly. Long-form and complex briefs go frontier. Short-form, structural and extraction work goes cheap. The middle is where the judgement is, and that's exactly where having actual numbers beats having an opinion.
Re-run the eval quarterly. Given four Flash models in 106 days from one vendor alone, a decision made in March is not a decision that holds in September. Search Engine Land and HubSpot are reasonable places to track how the wider industry is handling this.
FAQ
How do I evaluate AI models for marketing? Build a fixed set of 10 real briefs from your workflow, run them through each model, score blind on brief adherence, voice match, factual reliability, format compliance and editing time, then route tasks by where quality justifies price.
Why are public benchmarks bad for marketing decisions? They're vendor self-reported, saturated at the top, and measure reasoning capabilities with no established correlation to copywriting or brief adherence.
How many tasks should be in my eval set? Ten is the sweet spot: enough to cover task variety, small enough to re-run quarterly without it becoming a project.
Why does the test need to be blind? Because scorers who know which output came from the expensive model rate it higher. Blind conditions are the only way to get an honest signal.
What's the most important scoring dimension? Brief adherence, weighted around 30%. It's the most objectively scorable and the most predictive of real-world usability.
Why measure editing time? Because an editor's hour costs far more than a million output tokens. Editing time converts output quality into money, which is the only comparison that matters.
Can I test Claude Mythos? No. It's gated to vetted cybersecurity and life-sciences professionals. Test Fable or below.
Can I access GPT-6 Astra to test it? Maybe not yet. At rollout it was off by default for enterprise admins, check with whoever owns your workspace before scheduling the test.
How often should I re-run the evaluation? Quarterly. Model releases and price cuts in 2026 are fast enough that older decisions go stale quickly.
Does the eval tell me to pick one model? It shouldn't. The useful output is a routing table mapping task types to model tiers, not a single winner.
If you want help building an eval set that reflects how your team actually works, I'd be glad to talk. I've spent 4+ years in marketing helping edtech and startup brands grow organically. You can see the work and get in touch through the contact form at younusfardeen.com.