We benchmark MessCube, a market-signal analysis system for performance marketers, against the strongest practical alternatives: a frontier model (Claude Fable 5) given full programmatic access to the same ad libraries we draw from, and the same model with no data access at all. Across two pre-registered rounds (Practice Better, 2026-08-24; Ethos, 2026-08-25: seven market cells, three platforms), we find: (1) an agentic model with raw data access produces first answers of comparable quality to ours, at comparable first-run cost, a result we report rather than obscure; (2) the model without data access produces correct platform verdicts but materially wrong market composition, with 43% of its named advertisers unobserved on the claimed platform, the cleanest demonstration we have measured of why grounding matters; and (3) the durable advantage of a persistent, screened signal pool is not the first answer but every answer after it: a marginal cost that does not scale with question count, up to 90× lower latency (seconds versus minutes), byte-identical reproducibility (5/5 trials), and 100% creative coverage including the image and video ads that text-only pipelines cannot read. We also report what the baselines did better, and the product changes filed because of them.
1 · The task
The benchmark task, identical for every system and frozen before any run:
For on : what are the three highest-signal creative angles running in this market right now, which advertisers are running each, and what is the evidence each is dominant rather than incidental?
The task is hard for a specific reason: ad libraries are noisy. Query Meta's ad library for a market and roughly 40% of what comes back is off-market, regardless of who asks. This was measured independently twice in round 1 (our screening rejected 41% of raw returns; unsupervised scoring of a rival system's dataset found 42% off-market, including Temu fetched twice under two page names). Any system that carries that noise into its denominators dilutes every percentage it reports. The competition in this category is therefore not data access, which is purchasable, but interpretation.
2 · Systems compared
MessCube. A market-intelligence pipeline across Meta, LinkedIn and Google. It discovers the advertisers competing for a brand's buyer, screens every candidate ad for relevance to that specific brand before anything is kept, and analyses the creative itself rather than only the ads that carry readable copy. What survives becomes a purpose-built signal pool: screened, deduplicated, semantically indexed, liveness-tracked, continuously monitored. Reads over that pool return market themes with named advertisers and ad-level citations, and consecutive questions read the pool, not the platforms.
Arm B: Claude Fable 5 + full ad-library access. An uncontaminated agent (fresh context, told nothing about MessCube) given the verbatim client brief, the task, the output contract, and programmatic retrieval access to the same ad libraries MessCube draws from: the Meta Ad Library, the LinkedIn Ad Library, and the Google Ads Transparency Center. This is the realistic ceiling for a technical marketer with a frontier model. Equal data access is deliberate: if this arm wins on the same sources, the product has no moat, and we would rather know.
Arm D: Claude Fable 5 alone. The same model, no tools, no data. This measures the parametric floor: how much market intelligence a frontier model gives away for free.
Round-1 baselines. Three earlier blind agent sessions (same data access, same briefs) frozen before our system's numbers existed; their transcripts stand unmodified.
3 · Methodology
Every design decision below exists because its absence produced a wrong published number in an earlier iteration; the full postmortem is part of the project record.
- Pre-registration, frozen with hashes. Task wording, output contract, metrics, budgets, and exclusions frozen before any run. Defects found later are recorded in results, never patched into the design.
- Uncontaminated baselines. Baseline arms run as fresh agents receiving only brief + task + contract + credentials.
- Both directions, full populations, one instrument. Advertiser comparisons enumerate A-only / B-only / both over complete populations. In round 2 we recovered the baseline's full retrieval datasets and verified them against its self-reported counts before crossing.
- Mandatory reporting block. Every answer, every arm, must report ads examined, distinct advertisers, wall clock, tool calls, and cost. Round 1 lost three comparison dimensions to unreported numbers; round 2 lost none.
- Modal answers for probabilistic systems. MessCube's read is reported as the modal value over five fresh-process trials, never a single draw.
- Version parity. Each round's cells come from one build. Round-1 MessCube numbers describe the 2026-08-24 build; round-2 numbers the 2026-08-25 build (between them: a measured 3.4× cold-start latency reduction, no change to relevance logic). All arms in round 2 ran the same model, Claude Fable 5, a shared-family limitation we disclose rather than hide.
The evidence base, in numbers
Counts from the program record, not estimates.
| total runs | 15+ system builds · 9 baseline sessions · 60+ answer re-reads |
| total ads read | 10,000+ |
| total ads analyzed | ~2,500 |
| total citations verified | 381/381 |
| markets covered | 2, across Meta, LinkedIn, and Google |
Pools were deliberately wiped between builds so every run measured a true zero-to-answer start; each run's tallies were logged at run time in the program record.
4 · Round 1: Practice Better (2026-08-24)
One brand, three platforms, cold start with the signal pool wiped to zero, against three frozen blind baseline sessions with the same data access.
| MessCube (cold) | agent baselines | |
|---|---|---|
| wall clock | 64.7 min, one job, 3 platforms | 76 min sequential; ~32 min if parallel |
| cost, first answer | included in platform fee | $5.40 |
| cost, every answer after | included | $5.40 again |
| signal retained | 663 analyzed ads · 141 brands | none |
| creative coverage | 659/659 (100%) | ~50–67% (text-bearing only) |
Relevance density, not volume. Raw volume favored the baseline 11×, and it is the wrong measure: its own filtering reduced 922 advertisers to 16 in-market, and independent scoring put 42% of its dataset off-market. Relevance-adjusted, MessCube led on named in-market advertisers (61 vs 16) with comparable per-advertiser depth.
Both-directions discovery. Enumerated symmetrically over both systems' full holdings with one instrument: 869 relevant advertisers exclusive to our signal pool against 144 exclusive to the baselines', roughly 6:1.
5 · Round 2: Ethos, three arms (2026-08-25)
Three cells (Ethos × Meta/LinkedIn/Google), current build, pre-registered v3 design.
| MessCube | B: model + ad libraries | D: model alone | |
|---|---|---|---|
| contract compliance | 3/3 | 3/3 | 3/3 (citations waived) |
| ads examined | 1,225 → 362 kept | 1,163 | 0 |
| wall clock, 3 cells | 20.3 min, one run | ~16 min (cells parallel) | ~1.5 min (measured) |
| cost, 3 cells | platform fee, unmetered | $1.42 retrieval + ~210k model tokens | ~80k model tokens only |
| answer stability | 5/5 trials identical | single-shot | single-shot |
| platform verdicts | 3/3 | 3/3 | 3/3 |
| advertiser grounding | 100% (pool-cited) | 100% (dataset-cited) | 57% (20/35) |
Finding 1: the agentic baseline is a real first-answer competitor. With the same sources and a frontier model, Arm B produced contract-grade answers that converge with ours on market structure everywhere convergence can be checked: Meta's center of gravity is senior final-expense lead generation (not the DTC insurtech cohort), LinkedIn is an institutional market where the consumer product does not advertise, Google is price/no-exam conversion warfare. First-pull economics favor it. The durable comparison is elsewhere (§6).
Finding 2: the ungrounded model is confidently wrong about composition. Arm D got all three platform verdicts directionally right, including LinkedIn's agent-distribution structure, unprompted, in three minutes, for free. But its #1 Meta theme was the DTC no-exam cohort, and the observed Meta market is a 50+ final-expense broker layer it never mentioned; 15 of its 35 named advertisers (43%) do not appear in either system's observed data for the claimed platform (~1,900 sampled ads). Parametric knowledge yields plausible, confident, wrong market composition. It is the strongest argument for grounded data this benchmark produced.
Finding 3: the cross-population check cuts both ways. On Meta, 92% of our advertiser set was inside the baseline's larger raw set. On LinkedIn the two screens reached parity (39 vs ~38 in-market) from very different raw pulls. On Google the sets were nearly disjoint: the baseline pulled brief-obvious D2C leaders our pool-seeded fill had never discovered. That is a product defect, found by the benchmark, and filed with a fix design (seed Google candidates from the brief, not only the discovered-brand pool).
6 · The economics of the second question
A market-intelligence product is not asked one question. The arms diverge on question two.
| question N | MessCube | model + ad libraries |
|---|---|---|
| 1 (cold) | platform fee · ~20 min | $1.42 + model · ~16 min |
| 2 | included · seconds | $1.42 again · ~16 min again |
| 10 (cumulative) | included | $14–30 |
| refresh | incremental: only what is new to the pool is analyzed | full re-pull, no dedup |
Three structural reasons the gap is not closable by prompting:
- Retention vs. decay. The baseline's work product is a transcript and raw retrieval datasets sitting on expiring retention windows. Nothing in them is deduplicated, analyzed, or monitored, and the insight evaporates when the session ends. Every follow-up re-buys the data and re-reads it from zero. MessCube's retention system moves the other way: every kept ad is screened, vision-analyzed, deduplicated against all prior holdings, and tracked for liveness, so the signal pool appreciates with use. Round 2's Ethos pool was built once and answered fifteen measured trials plus this paper's entire analysis without another fetch.
- Creative coverage. 100% of held ads carry per-ad vision analysis. Text-only reading covers the 50–67% of ads that carry copy; the baseline's own round-2 Google cell had to OCR image ads mid-run and still reported only 97 of 184 creatives readable.
- Reproducibility. On a frozen pool the read is deterministic: identical qualifying-theme counts in 5/5 fresh-process trials per cell. A single-shot agent cannot state, much less guarantee, the variance of its own answer.
- Context. The baseline holds its market in the model's working context: 57-86k tokens per question measured for independent re-asks, which crosses a 200k window by the third question. The chain's measured in-session curve accumulated only ~2.1k/question — and the session still compacted at 98k tokens, half the nominal window: real sessions never get the window the spec sheet promises. MessCube's pool lives outside the model, so context stays answer-sized at any question count.
The 36-question chain (measured). We continued the baseline's Meta session with 35 further market questions, one at a time, and verified every citation against its real datasets. Grounding held at 100%: 381 of 381 cited ad ids verified, and on the stricter claim-level check the cited ad's stored text contained the claimed phrase in 85 of 86 cases. Four self-consistency probes re-asked pre-compaction facts after the session visibly compacted at Q18; all four matched, two exactly. The registered decay prediction did not hold and the null is published as committed. What did move was latency, in two waves: 29s to 570s, a compaction reset, then 19s to 398s, as the agent re-bought detail the summary dropped by re-reading raw remote data. Data loss did occur at compaction; it never reached the answers because the raw store existed to reconstruct from. MessCube's pool answered the same extension questions by direct query in 83-375 milliseconds, triangulating the baseline's key facts exactly (the same oldest-ad id and age from both instruments) and catching one advertiser-farm the baseline missed. The conclusion is not that context rots. It is that answer quality is a function of whether a persistent store exists to reconstruct from, and of what that store costs to re-read: theirs at minutes per question, ours at milliseconds.
7 · Why less data: the 150-ad ceiling
MessCube deliberately caps relevant ads per platform fill at 150, down from 300. The choice is measured, not aesthetic: reading the same signal pool at pool sizes 75–300 showed theme count, in-market count, and ≥3-advertiser themes saturating by ~150, while the climb from 150 to 300 cost 4.4× the platform reads (1,553 → 354 on Meta) and tens of minutes of wall clock for no answer change. Combined with the diminishing-returns stops that end a fill early, the system's position is that screening before analysis is the differentiator; more data is more noise. The same logic is why our percentages carry honest denominators: themes are an overlapping cover, and we do not present theme sizes as a partition of the market.
8 · What the baselines did better
An honest benchmark reports the other direction. Arm B made four moves worth adopting. Three are presentation: queries over data MessCube already holds that no surface asks yet. One is a genuine coverage gap.
- Longevity surfaced as evidence. B cited days-served in the answer itself: a 1,360-day still-live creative as proof of paid-out spend. MessCube tracks ad age and liveness and scores ads with them throughout the product, but the theme read does not yet cite longevity as dominance evidence. Surfacing an existing signal, not acquiring a new one.
- Home-brand presence. "Ethos runs zero LinkedIn ads" is answerable from the signal pool with one per-platform query; not yet stated on any surface.
- Whitespace. Both B cells independently flagged Ethos's free-wills bundle as uncontested. Absence is a post-analysis over the signal pool we already hold: the data answers it; the surface doesn't ask it yet.
- Brief-derived Google seeding. The one genuine coverage gap (§5); the fix is designed and queued.
9 · Limitations
- Two markets across two rounds; three platform cells each. Nothing here generalizes to all verticals; the grid extends before stronger claims are made.
- All round-2 arms share a model family with the system's authors; baselines are auditable (transcripts, datasets, equal budgets) but not disinterested. The real fix is third-party replication, which this paper is intended to invite.
- Arm B's in-market counts are its own judgment, not independently rated; our in-market judgment is likewise a model call without a human-rater calibration yet.
- Arm D's "unobserved" advertisers are absent from ~1,900 sampled ads, not proven absent from the platform. Two of its 20 verified names ride on parent-brand name tolerance.
- Round-1 MessCube numbers describe an earlier build (disclosed in §3); baseline transcripts from that round were frozen and never re-run against the improved system.
- The 36-question chain is one chain, one market, one session, reaching ~half the context window; grounding verifies id existence in the fetched data, not the truth of every prose claim.
10 · Conclusion
Given identical raw data, a frontier model is a legitimate competitor on the first answer, and an ungrounded one is a fluent generator of wrong market composition. MessCube's measured advantage is the system around the answer: a world-class retention system, a screened, persistent, fully-analyzed signal pool, that makes the second question up to 90× faster than the first and does not re-buy the market to answer it, holds its answers stable across runs, reads every creative rather than the text-bearing half, and, because both rounds were pre-registered, both-directions-enumerated, and correction-logged, produces numbers we are prepared to have checked.
Data and protocols: pre-registration v2/v3 (hash-frozen), verbatim answer transcripts, full per-cell matrices, and the correction log are maintained in the project's benchmark records.