The Go-to-Market Context Benchmark

In go-to-market work, the architecture you choose for delivering context — not the amount of context — is what determines output quality. The best outputs don't come from the most tokens. We ran the same tasks through six context architectures and graded the output blind to show it.

Once you've accepted that your team needs a context system, the next question is its shape — and unlike most of the AI-tooling conversation, this one has data behind it. We built six ways of giving an AI model the same go-to-market knowledge, ran nine real GTM tasks through each, and had a blinded judge score every output against a weighted rubric. Same knowledge, six containers. The container is the variable.

What we tested

Six architectures, holding the knowledge constant and changing only its shape: a no-context control; a single inlined document; a router over a folder of files; a navigable wiki; a typed knowledge graph; and a layered system (graph plus narrative). Nine tasks — account plans, business cases, call prep, objection handling, campaign briefs. Each output graded blind on a weighted rubric, with cost modeled from the token traces.

Architecture beats volume

The cleanest finding in the set: the same knowledge in a navigable wiki cost more and scored lower than in a typed knowledge graph. Same facts, worse container — more tokens burned, more navigation turns, a lower score at the end. Token spend ranged roughly 30× across the field, and past a competent baseline it was negatively correlated with quality. The elaborate, expensive architectures were frequently the worst ones.

Two jobs, not two competitors

The result that matters most for a buyer resolves a false tension. The typed knowledge graph — the canonical layer — is what you run your go-to-market on: it was the efficient frontier, top-tier quality at the lowest cost, and it never lost a task. The layered system — graph plus a narrative layer — was the quality ceiling, and its edge showed up exactly where the work was deep and conceptual: synthesis, positioning, the strategic reframe. Those are two different jobs. One is the daily operational workhorse; the other is what you reach for when the thinking is hard. The data shows each winning its own job — not one beating the other.

This is why the architecture choice is a fit, not a leaderboard. The right shape depends on what your team actually does most.

Why a 70 isn't an 80 — show the work

A fair challenge: if the floor scores 70 and the winner scores 80, is that a real gap? It is — and seeing why requires showing the grading, not just the number. A no-context model still writes a structurally complete, plausible plan and banks partial credit on the easy criteria. So 70 means "looks like a real plan, misses what wins the deal." The whole spread lives in a handful of decisive, high-weight criteria — the deal's most time-sensitive fact, the one named comparable the CFO asked for — where a competent-looking output scores a flat zero because it never mentions the thing that matters. The average is honest; it just dilutes a decisive miss across a dozen criteria where everyone looks fine. We publish the full rubric, the blinded scores, and the criteria so you can audit it.

What this proves — and what it can't

Read the altitude carefully. This benchmark proves the architecture choice is consequential — that how you shape context has real, measured effects on quality and cost. It does not prove you need a context system in the first place; that's an organizational argument, made in the piece before this one. And the honesty that makes an open benchmark worth trusting: this is a first cut, single-rep, one author's gold standards. The patterns are directional, the methodology is fully published, and more replications are the public roadmap. We'd rather show you the seams than hide them.