Every marketing team is being told to point AI at its market. Almost nobody has measured what happens when you do. So we did, against ourselves, with the deck stacked in the AI's favour, and we are publishing the parts we lost.
The test
Two real brands. One question each, on Meta, LinkedIn and Google: which creative angles are winning in this market right now, who is running them, and what is the evidence?
Three systems answered it. MessCube. A frontier AI, Claude Fable 5, given the same access to the same ad libraries we use, on purpose, because if it wins on equal data we have no product and would rather know now. And the same AI with no data at all, answering from what it already knows, which is what most "AI for marketing" is under the hood.
We wrote down what we expected before we ran anything, published every answer verbatim, and checked every claim against the ads themselves. Six market cells, over ten thousand ads read.
Finding one: on the first question, the AI tied us. And cost less.
Given the same data, the frontier AI produced a real, well-evidenced answer. It found the same market structure we did. On that first question it was cheaper to run than we were.
If your team asks one question about your market per quarter, buy the AI. We mean that.
Nobody asks one question. You ask what is working, then who is running it, then what changed since last month, then what the new entrant is doing, then all of it again for the other platform.
Finding two: every question after the first
The AI has no memory of your market. Every question re-buys the data and re-reads it from scratch, so the second question costs exactly what the first did, and so does the fiftieth. MessCube keeps a screened, analysed store of your market, and the second question is a lookup.
We ran it to the end: 36 consecutive questions in one session. The AI's answers stayed accurate. We had predicted they would degrade, they did not, and we published that. What degraded was time. By question 18 one answer took 9½ minutes. The same questions against our store came back in under half a second.
Multiply that by what your team costs per hour and the cheaper option stops being cheaper somewhere around question three.
Finding three: AI without your market's data invents it
The third contender, the AI with no data, was fluent, confident and specific. It named competitors. It ranked angles. It sounded exactly like the other two.
43% of the competitors it named were not running ads where it said they were. Real companies, wrong market, stated with total confidence. Its top-ranked angle for one market did not exist there: it described the well-known challenger brands, while the money was going to a tier of advertiser it never mentioned.
It got the strategy right and the market wrong. That is the most expensive kind of wrong, because it reads as insight.
What this changes
The cheapest line on the invoice is the data. Fifty questions on the agent is $71 of it, and roughly thirteen hours of a paid person waiting for answers. Performance teams already burn 30 to 50% of creative budget testing concepts the market has ruled out, not through incompetence but through not knowing what has already been proven, and the waiting is how that not-knowing gets paid for. One option is a fixed fee against an asset that compounds with every question asked of it. The other re-buys the same market every time someone is curious.
What the asset changes on the ground is the picture itself. In our test, the AI's idea of a market was the brands you would name from memory. The ads actually running showed a different tier of advertiser doing the volume, with different angles, different price points, different hooks. About half of those ads were image and video, which a text-only tool cannot read at all. We analyse every one of them.
Every theme we report carries named advertisers and the ads behind it, so "how do we know?" has an answer other than "the AI said so." We checked 381 of 381 citations in our own benchmark against the underlying ads. And the read holds still: ask the same question five times and you get the same answer five times. A chat tool gives you a different one each time, which is fine until a quarter's plan is built on one of them.
The AI did four things better than us. It cited how long an ad had been running as proof of spend. It noticed a brand ran zero ads on a platform they had assumed they were on. It spotted an offer nobody in the market was making. And on one platform it found competitors our system had not discovered.
Three of those were things our data could already answer and our product was not showing. One was a real gap. All four are fixed or scheduled, and they sit in the technical paper with the same prominence as everything we won.
We publish the losses because a benchmark where the author wins everything is not a benchmark. It is a brochure. You can check ours.
See your market's actual read
Give us your brand. We will show you the angles running against you right now, who is running them, and the ads behind every claim.
Every number here comes from the full technical paper, published with its methodology, its limitations and the raw transcripts: the technical whitepaper.
