Fresh Context
Research · Fresh Context

The Jagged GTM Frontier

Why AI underdelivers in go-to-market and what to do about it

Read the news and AGI seems just over the horizon. In software engineering, data science, customer support, and cybersecurity, AI is already changing how knowledge work gets done.

But progress among GTM teams is far more uneven. Deploying the same models, most GTM teams aren't seeing AI transform their function the same way. In 2026, 31% of chief sales officers named proving the ROI of AI tools among their top challenges. Gartner

Why is that?

DomainCapabilityThe jagged frontier(Domain)ChemistrySolve protein foldingCybersecurityIdentify software vulnerabilitiesGo-to-marketWrite a prospecting email
AI Capability by Domain
Each ray is a domain; each dot a real task, placed by how well AI does it.

In some domains, AI has demonstrated capabilities at or beyond human limits.

Solve protein folding

The 2024 Nobel Prize in Chemistry recognized Demis Hassabis and John Jumper for AlphaFold, which cracked a fifty-year challenge in protein-structure prediction and helped researchers predict the structures of nearly all known proteins. Structure work that once took a lab years now starts from a lookup.

Nobel Prize 2024
Identify software vulnerabilities

In the first month of Project Glasswing, Anthropic and roughly 50 partners used Claude Mythos Preview to find more than 10,000 high or critical-severity vulnerabilities across widely used software systems.

Project Glasswing
Write a prospecting email

Ask a frontier model to write a prospecting email and the default result is stilted, slightly off, or completely wrong for your business. Why does AI excel at one task and stumble on another that appears much simpler?

The jagged frontier

AI capability is jagged: sharp and reliable on some tasks, unreliable on others. Two tasks that look equally hard to a person can land on opposite sides of that line - one the AI nails, the other it fumbles. Researchers call the line the jagged frontier. The model succeeds or falls based on its training data, what it can see at runtime, how the work is specified, and how success is defined.

Dell'Acqua, Mollick, et al.

GTM is a messy, tribal, team sport. It's always taken a combination of relationships, skills, deep experience and functional expertise to bring complex products to market. Despite the challenges, there are AI-native GTM teams driving far more revenue per employee and reinventing the way they engage the market with AI as the foundation.

What can we learn from them?

We aren't folding proteins in GTM because leaders are still trapped in the Predictable Revenue mindset - the new frontier is finding the value in your context and providing it to your next best buyer.
Jordan Crawford, stipple portraitJordan CrawfordFounder, Blueprint GTM

Their edge isn't the models - everyone rents the same models. It's three things they've built:

  1. I.Context availability - store your knowledge so AI can use it.
  2. II.Task clarity - define the work so AI knows what done means.
  3. III.A coordinated way of working - so one person's win becomes the whole team's.

The first two make a single AI request work. The third makes the wins compound across a team. No one was born AI-native - they built this, and so can you.

Why is AI so good at software engineering?

Drop a language model into a codebase and it can start building features and tackling bugs with very little coaxing. Product managers, engineers, and developers may have more technical experience and fluency than GTM teams, but the main reasons AI is more impactful in their domains are structural.

Go-to-MarketAI is default incapableSoftware EngineeringAI is default capableanalyticslegalTask Claritydoes AI know what done means?Context Availabilitycan AI see what the work needs?LowHighHigh

Context Availability

Engineering gets this for free: the codebase.

When an LLM has access to a codebase, change log and docs, the human operator doesn't need to explain the operating domain. The AI can read it. Context availability is extremely high for engineering teams, but fragmented by default for GTM orgs. The more complex the business, the more your GTM context is scattered across shared drives, Slack messages, LMSs, call libraries, and tribal knowledge.

Task Clarity

Engineering gets this for free too: the spec and the tests.

It's not enough that AI can read what already exists. It needs a clear signal for when its work is done. Engineering has that signal built in: a feature ships against a spec and the test either passes or it doesn't. GTM doesn't work that way. "Done" for a homepage layout or an email subject line is a judgment call. Nobody signs off on "most likely to get opened" the way a CI pipeline signs off on a pass. That's the gap: AI can execute the GTM task, but usually no one has defined what counts as finished.

GTM gets neither by default.
Here's how AI-native teams build both.

You can spend years arguing ontologies, classifying old docs, and trying to test a perfect GTM context system - you'll never get there. Like everything else in GTM, context is only as good as the outcomes it drives.

Don't boil the ocean. Take a specific moment in time - a product launch, a new campaign, an initiative launching at SKO - and start building your context system to support the GTM work that reaches your buyer. Train your team to treat context as a tool to hit the number, and keep it living: fold in what the market hands you the day it lands, prune what it moved past, and check anything that goes stale the moment you use it. That's what fresh means - not a release calendar, a daily practice.

We didn't arrive at any of this by reading. For three years our job was putting AI inside the daily work of a go-to-market team, and most of what we know here we learned by building the wrong thing first. We hand-rolled retrieval before we understood what a canonical layer was for. We built Rube Goldberg orchestrations when what the team needed was one good skill. We stood up an account-research bot that produced beautiful piles of text nobody read, then threw it out for scoring tuned to the few signals that actually predicted a good account. Every condition named on this page has a version of us getting it wrong underneath it.

That practice is also where this article came from. Below is a living slice of the context system it was written against - the working graph, not an export. Start from this article's own node and walk outward to see what it drew on.

One thing this is not: the whole of context, or the pattern for all of it. What you are about to browse is a conceptual system - arguments, definitions, evidence, the thinking. Your CRM is context too. So are your product data, your call recordings, your pricing, your win/loss notes. Those want different systems with different owners and different refresh rates, and a GTM team getting real leverage from AI will run several of them. One architecture, many systems. This is the one this page came out of, and it is the easiest one to start.

It's still growing. On Monday, Colin Fleming - the CMO running OpenAI's business marketing - published What happens when everyone can do marketing? It lands on two arguments this page makes - hand-offs are where ideas die, and the leader's job is setting the standard for what deserves to ship - so it went straight into the graph. Look for the newest node. Fleming, LinkedIn

For years, marketers have pursued the golden record, a unified view of who the customer is. In the AI era, we need Golden Context. Not just who the customer is, but what they need now, what the business wants to accomplish, and what our systems are actually capable of doing about it.
Scott Brinker, stipple portraitScott BrinkerAnalyst & advisor, chiefmartec

Context availability is everything the AI can see and reason against as it works. Without context from inside your walls, a model can only regurgitate public knowledge or make things up. And it's not one thing: some context is built ahead of the work - positioning, voice, the claims you stand behind - and some gets fetched in the moment from the systems your team already runs. Both kinds need curation. Good GTM context is your sharpest thinking, not your biggest export.

The harder discipline is keeping it current. Last quarter's positioning ages fast, and context only stays an advantage while someone owns it - pruning what the market moved past, updating what the product changed, versioning what the team relies on. Context nobody maintains decays back into the noise it came from.

The source of alpha is the context your business can uniquely assemble - kept current, and put into production. That's when AI starts transforming GTM instead of simply accelerating the old motion.
Zach Vidibor, stipple portraitZach VidiborCo-founder & CEO, Octave

In GTM, change is constant and context is king.
AI-native teams build living context systems.

Good AI works against the same specific motion you do - the job on every task is specific to that deal, campaign, channel trend, and day of the week.

Anything you don’t provide about the task, the LLM assumes. AI needs to know everything about the job you’re giving it, but can only handle so much text before output degrades and your token costs balloon. Create repeatable, composable, and shared skills and data sources to get better AI output on every request, and at a lower cost.

  1. You can write mega prompts and dump context in for a one-off task, but if you want AI to provide repeatable value, you have to give it repeatable task frameworks. Here’s a cold email prompt to the same prospect, composed three ways - watch two things about each result: did this run work, and would the same approach work the hundredth time someone else on your team runs it?

    The job, written into one prompt.

    Approach 1 · Everything in the promptprompt
    Write a cold prospecting email to Dana Reyes, VP of Sales at
    Northwind, a mid-market SaaS company. We sell a revenue intelligence
    platform. Keep it short, mention we help teams like theirs, and ask
    for 15 minutes next week.
    response
    Subject: Quick question, Dana
    
    Hi Dana, I hope this email finds you well! I'm reaching out because we
    help mid-market SaaS companies like Northwind unlock powerful revenue
    synergies and drive efficiencies across the sales org. We've built an
    innovative platform that I'd love to show you. Do you have 15 minutes
    next week? Thursday works great on my end!
    this run:failsacross the team, over time:fails

    Too little to go on, so the model guesses - and every rep guesses differently. The work was never really defined. This one is a genuine individual-execution problem.

    A sparse prompt fails.

    You can try fixing it by putting everything into the prompt.

    Approach 2 · Everything hardcodedprompt
    You are our SDR. Write a cold email to Dana Reyes, VP Sales at Northwind.
    
    # Who they are
    Mid-market SaaS, 200-2000 employees, Salesforce shop, struggling with
    forecast accuracy and rep ramp. [ + 38 lines pasted from the ICP doc ]
    
    # Voice
    Be bold, punchy, confident. Challenge the status quo.
    (from the 2024 brand deck: be understated, consultative, never overpromise.)
    
    # Product
    Revenue intelligence platform: call scoring, deal risk, forecast rollups,
    pipeline hygiene, and Momentum. (Momentum was sunset in Q2 — do not
    mention it — but the winning examples below still use it.)
    
    # Rules
    Never sound salesy. Always create urgency. Keep it under 120 words, but
    include two proof points and a customer story.
    
    # Winning examples — copy the structure, be original
    [ 3 full emails pasted in ]
    response
    Subject: Northwind's forecast is probably fiction
    
    Hi Dana — bold claim: most SaaS forecasts are fiction. That said, we'd
    love to explore whether there's a fit. Our Momentum module helps teams
    like yours get ahead of deal risk, and we've seen real results. Open to
    a quick chat? No pressure — happy to send a one-pager first.
    this run:sometimesacross the team, over time:fails

    The same prompt might or might not work on a different campaign, account, etc. Six months later, you’ve forgotten what went into your mega prompt. It’s a pain to update. There are new features, campaigns, and competitive positioning, but your AI is stuck in the past. Hardcoded prompts propagate, drift, and eventually become context debt. The more context debt you carry, the harder it is to sustain GTM acceleration.

    Build prompts from live, fresh context.

    Give clear task objectives and intelligently pull the live GTM instructions, templates, and data you need to get the best results from every AI prompt, and save tokens.

    Approach 3 · A composable specspec
    Cold email to Dana Reyes, VP Sales at Northwind.
    Goal:  book 15 minutes.
    Shape: 4 lines — their trigger, one proof point, a soft ask.
    
    Reference our shared definitions ↓
    composing · looking up the current definition from each system
    value-prop(“Northwind-full platform, new logo midmarket”) lead with forecast rollups + deal risk · Momentum sunset Q2 (omit)
    brand-format(“prospecting: above the line”) understated · one proof point, not three · no “unlock synergies”
    buyer-stage(“cold”) name the trigger · one proof · a soft 15-minute ask

    Define these once. This email, a teammate’s, and next quarter’s all look up the same systems - update one and every prompt gets the new answer.

    response
    Subject: Northwind's Q3 forecast slip
    
    Hi Dana — you added ~40 reps this year, and forecast accuracy usually
    breaks right about there. We helped Evergreen cut forecast error 32% in
    a quarter. Worth 15 minutes Thursday to see if it maps to Northwind?
    this run:worksacross the team, over time:works

    The spec stays short because the parts that don’t change - who it’s for, how we sound, what’s true, the play - are defined once and looked up, not rewritten. Define the work well once, and every prompt reuses it.


    This isn’t about wording one prompt well. Repeatable AI output comes from treating your task definitions as shared assets - each with a single owner, kept current as the business moves - so every prompt draws on one source of truth instead of re-explaining the job from scratch. That’s a discipline your organization takes on, not something the agent does in a single turn.

    In software there's build time and runtime. Build-time context you curate ahead of time - canonical, up to date, something you can point to at a moment in time. Runtime context you go fetch when you need it - the CRM, Snowflake - and assemble on the fly.
    Jacob Dietle, stipple portraitJacob DietleFounder, Taste Systems
  2. Every wrong draft is a fork. You can edit the output - patch this one and ship it - or iterate upstream: fix the system that generated it, then regenerate. The first is faster today. Only the second compounds.

    Downstream editing
    generateoutput✎ fix the outputship
    The fix ships once, then it’s gone. The next task starts from scratch.
    Upstream iteration
    ✎ fix the systemgenerateoutputship
    The fix lives in the system, so every output after it starts higher.

    A downstream edit dies with the output. An upstream fix - a sharper definition, a corrected rule, a better prompt - lives in the system, so every output after it starts higher. Empowering the team to improve the system as they work drives adoption and acceleration.

  3. Taste feels unmeasurable - until you name what “good” means as weighted criteria and score every run against it. It also clears the bottleneck AI created: when drafts arrive in minutes, approval becomes the constraint - six review rounds for a deliverable an agent produced in ninety seconds. Two jobs: tune the system until it passes, then keep it passing as the inputs change.

    Tuning

    Run until the scores go green.

    runs
    01
    02
    03
    04
    05
    06
    weight
    weighted score0-100, weighted by the criteria below
    40
    51
    65
    77
    86
    93
    100%
    Trigger fitNames a real reason to reach out now
    42
    55
    64
    76
    85
    93
    25%
    Proof pointOne concrete, credible result
    30
    45
    62
    72
    84
    90
    20%
    VoiceSounds like us, not a generic bot
    38
    40
    58
    74
    82
    94
    20%
    ConcisionTight - no filler, no throat-clearing
    55
    60
    68
    78
    86
    90
    15%
    Clear askOne specific, easy next step
    60
    65
    75
    85
    92
    98
    10%
    No stale claimNothing sunset or out of date
    20
    50
    70
    85
    95
    100
    10%
    Early runs score red. Each run tightens the system - criteria, definitions, checks - and the scores climb, left to right, until they’re green.
    worsebetter

    Maintenance

    Every new input, the same bar.

    variants
    A
    B
    C
    D
    E
    F
    weight
    weighted score0-100, weighted by the criteria below
    95
    92
    94
    93
    93
    92
    100%
    Trigger fitNames a real reason to reach out now
    95
    94
    96
    90
    95
    97
    25%
    Proof pointOne concrete, credible result
    92
    96
    93
    88
    95
    91
    20%
    VoiceSounds like us, not a generic bot
    96
    76
    95
    94
    96
    95
    20%
    ConcisionTight - no filler, no throat-clearing
    94
    95
    92
    96
    79
    95
    15%
    Clear askOne specific, easy next step
    97
    96
    84
    95
    98
    96
    10%
    No stale claimNothing sunset or out of date
    98
    99
    97
    100
    96
    68
    10%
    Green doesn’t mean done. Every new input - a different account, a fresh segment - is scored against the same criteria. Mostly green, with the odd dip you catch and fix. That is the standard holding.
    worsebetter

    Do that, and taste is no longer a gut call - it’s a standard the whole team measures against, every time. We’ve used a prospecting email above as the example, but AI isn’t just text generation. The same method - define the work, iterate upstream, score against a bar - builds any go-to-market task:

    • Landing page
    • Campaign design
    • Competitive battlecard
    • Customer emails
    • Event operations
    • Pipeline forecast

    Identify the leverage points where AI can provide the most value and engineer the systems to deliver it.

Re-architecting a GTM stack around AI isn't about delivering enough point solutions or one-off workflows, it's a product job: someone defines what good looks like, owns the system, and is accountable for the outcome. Before you buy more tools or tokens, align on a vision and a mandate to change.
Mike Rizzo, stipple portraitMike RizzoFounder & CEO, MarketingOps.com

Across any GTM task, applying AI is judgment and elbow grease.Across a GTM org, applying AI takes some magic.

AI-native GTM teams commit to a coordinated way of working.

Everyone in GTM has a different role to play, every role has different tools and ways of working. Giving everyone AI models and tools is a starting point - your GTM engineers, sales hackers, and ops talent will self-teach, acquire AI superpowers, and supply most of the real intelligence about what works. But even with the pull of brilliant early adopters, GTM teams still execute largely the same work they did pre-AI. Unless GTM leadership proactively redesigns functions, processes, and how people work together, the learning stays local and the benefits stay limited.

GTM AI TransformationIndividuals use AI to do theirpre-AI job slightly fasterLeaders redesign workaround AI capabilities

Redraw Roles

So much of GTM execution is hand-offs based on knowledge or functional expertise. Hand-offs are slow and expensive. AI is already extremely capable at functional tasks that previously required an expert - reporting and campaign ops, web development, deep market and account research, content synthesis. Every pre-AI GTM hand-off - for knowledge or functional execution - should be challenged. When a machine takes the drafting, the research, and the first pass, the work left for a person is what was always most important: judgment. A seat built to produce volume and a seat built to exercise judgment are not the same, and nobody gets from one to the other by being handed a faster tool. Leaders have to redraw the seats. AI-native GTM teams deliver more revenue per employee because they build the revenue machine to minimize hand-offs.

Standard AI Harnesses

The 'Harness' is the wrapper and worksurface around your AI model - it could be chat, cowork, or code. It's the set of skills and knowledge available to your team for every AI task. AI-native GTM teams design harnesses specific for their team's workflow and their business data protection and privacy needs, treat the harness as an internal product, and invest in its development and governance.

Shared GTM Context

With or without AI, GTM Context is fragmented by default. We invest in kick-offs, deal reviews and big launches because the biggest wins in GTM come from coordinated execution. Getting everyone together and aligned - establishing shared context - has always been expensive. With AI, people can build shared Claude projects, detailed prompts, skills or plugins, or maintain a shared context filesystem. When GTM Leadership designs context as a system to drive perpetual alignment, AI dramatically improves the speed and quality of execution.

The gap between theory and practice in GTM is leadership. Understanding what needs to be done, creating a shared vision, and coordinating work across fragmented teams has always been critical to GTM execution. In that way, rebuilding an AI-native GTM team calls on the same skills that earned GTM leaders their positions.

But AI is different. This change is massive and the frontier is jagged. You need to know how, when, and where to deploy AI, and every one of those is nuanced. AI's capabilities are continually changing. Frontier model progress continues to push the market to reimagine roles, rebuild software, and redesign processes. You could spend every day trying to keep up and every year rebuilding your GTM machinery, but neither delivers this quarter's number.

Most teams are trying to solve an org design problem with a tool purchase. We stopped asking which AI to buy and started asking how we want our team to work - and what it needs to know to work that way. Everything else follows from that.
Gurdeep Dhillon, stipple portraitGurdeep DhillonVP, B2B Marketing, Intuit

AI is more than capable of GTM execution -
once leaders and practitioners build the conditions it runs on.

Great Change is Great Opportunity.

Botanical illustration: an orange cut in two, the near half square to the viewer showing its segmented flesh, the far half turned behind it, with a citrus branch, leaves and a single orange blossom above, in full color

Let's acknowledge that AI didn't just change how we work - it changed the financial environment we're selling into. AI reset every company's product market fit. AI is helping competitors ship features and expand into new markets faster, and buyers are overwhelmed working their own AI adoption.

The GTM leaders excelling in this environment set the direction and the pace. They do the work to understand what AI can and can't deliver, use good judgment about where to deploy it, and define new ways for their teams to work. The job may be harder than ever, but the opportunity to make a difference is proportional.

This is a time for leadership. The leaders who design a GTM system that defines the way the whole team works with AI and deploys it on purpose are thriving. Point solutions, pilots, and Claude-pilled individual contributors can make a significant impact with AI, but only leaders can change the way that GTM is built. You can't have a coordinated way of working without someone with the vision and authority to coordinate it - at the leadership level and at the technical level.

The models won't be your edge - your competitors rent the same ones. The context you assemble, the definitions you set, and the way of working you build are yours alone. That's where the advantage compounds.

We're grateful to the operators and researchers who gave us their time, their experience, and their disagreement.

AI offers GTM leaders more leverage than we've ever had.
There's never been a better time to build.

Published by Fresh Context. We work at the front of AI and go-to-market every day - with clients, peers, and everyone else figuring it out. This page was built inside the systems it describes, and dated the day we believed it.

Sam Gong, stipple portraitSam GongFounder, Fresh Context · Author of this piecePublished August 2026

Contributed comments

Operators and thinkers changing the practice this research describes.

Across all the companies I'm working with, even the ones investing in go-to-market engineering, AI is siloed all over the place. A team might have some shared context, but the company does not. Until leadership fixes that, every win stays local.
Jared Waxman, stipple portraitJared WaxmanFounder, GTM Engineer School
Codify and evolve tasks just as you would a business process: what systems need to be interacted with, what decisions need surfacing, what should the outcome be? Build ways to incorporate v1 learning into v3, v4, and v5, and the whole org benefits from continuous low-cost transformation.
Shy Rahnama, stipple portraitShy RahnamaCo-Founder & Partner, AwareGTM

Acknowledgments

The jagged-frontier frame is Fabrizio Dell'Acqua, Ethan Mollick, and their co-authors. The technology-versus-organization gap is Scott Brinker's Martec's Law. The harness framing draws on Kyle Norton's picture of the GTM harness. The operators and researchers in Sources are the reason this page has evidence under it.

Sources

Dell'Acqua, Mollick, et al. "Navigating the Jagged Technological Frontier." HBS 24-013, 2023; Organization Science, 2026.

Nobel Prize in Chemistry 2024. Hassabis & Jumper, AlphaFold.

Project Glasswing. Anthropic, 2026 - 10,000+ high/critical vulnerabilities in month one.

Gartner. 31% of chief sales officers cited difficulty proving AI-tool ROI, 2026.

Kyle Norton. GTMnow, SaaStr.

The GTM Harness. GTM Strategist, 2026 - the harness metaphor and Owner.com's GitHub practice.

Reference deployments

Public examples of teams operating this way, each credited to the operator or lab whose published work documents it. None are Fresh Context engagements.

Owner.com - 3x revenue per AE. Kyle Norton's team runs one shared harness, versioned in GitHub. GTMnow

incident.io - five people. Tom Wentworth runs a full enterprise marketing stack on a five-person team.

Cursor - ChatGTM. Shared context plus a fleet of agents, reps authoring their own skills in natural language. The Signal

OpenAI - GTM on its own models. A Slack assistant for account context and meeting prep; the best model access still needed a context layer. GTM Assistant

Ramp - workflows as systems. Three-tier outbound with AI-generated email at tier two, and an internal Gmail overlay that classifies and prioritizes rep inboxes. Outbound Kitchen

Clay - agents in production. Claygent research agents and inbox-as-a-source, inside Clay tables. Changelog