How to evaluate a
pre-launch testing method.

Twelve questions to ask any vendor, including us. Tick them off as you go, or print the list and take it into the call.

Things to check
4

A measurement has to do four things. Anything less is an opinion with a number on it.

Questions to ask
12

Three per property. Ask them of every vendor on your list.

Rows we lose
3

Speed, writing and magnitude. Our answers are on the list too.

The four

Four things a measurement has to do.

Fail the first one and you cannot check the other three, because you have nothing stable to check against.

01

Repeats

Same page in, same answer out.

02

Shows its source

Every quote traced to a real page.

03

Reads like your buyer

Scans, hunts, abandons. Never patient.

04

Remembers

Week 4 knows what week 1 did not.

The checklist

Twelve questions. Ask any tool, including us.

Tick each one when a vendor answers it properly. Our own answers sit underneath in purple.

0 of 12 answered
01

Does it repeat?

This decides everything, and it is the cheapest to answer.

Good answer: yes, and they will show you.

Ours: yes. A re-run replays sealed findings instead of recomputing them.

Good answer: only the part you edited moves.

Ours: untouched findings come back identical, so the change is attributable.

Good answer: no. Verdicts are frozen under the version that made them.

Ours: if our environment drifts we say so and hide the comparison.

02

Does it show its source?

A made-up quote and a real quote look the same on screen.

Good answer: on demand, for any quote you pick.

Ours: quotes are checked against the fetched page, not taken on trust.

Good answer: every source carries an author class.

Ours: we learned this the hard way, then built the classifier.

Good answer: yes, marked as such rather than quietly dropped.

Ours: low confidence comes back labelled low confidence.

03

Does it read like your buyer?

A general AI assistant is the most patient reader alive, and none of your buyers are.

Good answer: a mix, weighted to the channel you buy.

Ours: cold LinkedIn runs 82% distracted, 15% skeptical, 3% shopping.

Good answer: readers leave, and it can say where.

Ours: every agent scans, hunts or abandons instead of reading to the end.

Good answer: committees do not average. One veto ends it.

Ours: roles combine by veto rules in code, not by blending scores.

04

Does it remember?

This separates a tool you use once from something you run a programme on.

Good answer: yes, matched by meaning rather than wording.

Ours: on four weeks of real evidence, zero findings carried over until we fixed it.

Good answer: a number you can find without asking.

Ours: 473 claims graded, 84% accurate, misses left in.

Good answer: yes. Grading after you know the result is not grading.

Ours: sealed before launch, then graded by your team against your own data.

Start here

Q1 costs four minutes and nothing else.

Do it before you take a single vendor call.

The four-minute test

Open two fresh sessions with no shared history. Paste the same page into both and ask the same question.

What it tells you

If the answers disagree about what is broken, you have an opinion generator. Useful, but not something to spend budget against.

Our answers

Three rows we do not win.

A guide whose author wins everything is a brochure.

We lose

Speed

An assistant answers in seconds. We take about thirty minutes.

We lose

Writing

We do not write copy. Drafting and variants are what an assistant is great at.

We lose

Magnitude

We tell you whether a call was right, not by how much. A live A/B test wins there.

Common questions

Answers, briefly.

Can I use a general AI assistant to test a landing page before launch?

For drafting and variants, yes. For a spend decision it fails Q1, because two fresh sessions give two different answers.

What if we connect the assistant to our own data?

It helps, and attaching evidence is easy now. It still cannot tell you it read the same thing last week.

Why does repeating matter so much?

Attribution depends on it. If the answer moves on an unchanged page, it says nothing about a changed one.

Is a high hit rate enough?

No. Ask whether the misses are counted, whether claims were sealed first, and who did the grading.

Before you spend

Fail in the test.
Win in the market.

Send us three campaigns you already ran and know the results of. We predict the ranking. You check it.