After 18+ months running AI search campaigns across now 40+ B2B SaaS companies, the biggest mistake we see isn't a missing piece. It's sequencing.
Some companies tweak schema and publish llms.txt files before they have a single page an AI would want to cite. Others chase third-party mentions before the core solution-aware pages even exist.
And plenty pour budget into top-of-funnel content that AI assistants now answer completely, so the buyer never sees the brand.
So in this article, I'll share the full framework we use at Rock the Rankings: the GEO Stack.
Three layers, built in a specific order, because each layer feeds the one after it.
And this time I'm going to get into the actual mechanics. How we build the prompt list. What the listicle looks like on the page. What the outreach email says. What the reporting looks like when a program works, with real numbers from a real engagement.
Before you start
A quick gut check.
The GEO Stack is the architecture. But architecture built on sand still falls over, and AI search will surface weak foundations painfully fast. Before this framework receives any investment, three conditions must be met:
01
A product that survives honest comparison
AI assistants compare mercilessly, and they don't always repeat your positioning spin. If your product genuinely loses head-to-head against the incumbents on the criteria buyers care about, GEO will put that weakness in front of more buyers, faster.
02
Willingness to make on-page changes and ship content
The methodology requires new pages, structural edits, and technical changes, at a real cadence. If those get stuck in a six-week approval queue, results stall with them. More on this at the end, because it's the number one way these programs die.
03
Third-party presence for the models to cite
AI assistants lean on review sites, editorial roundups, and community discussion. If your G2 profile is empty and nobody has ever written about you, Layer 2 starts from a colder position. You need to be ready to execute on the outreach layer, the community layer and the review collection layer as well.
The framework
The three layers.
Every GEO engagement we run uses the same three-layer architecture. The emphasis shifts by client and by month. The layers themselves don't change.
Solution-aware presence
Listicles, alternatives pages, comparison pages, use case pages. Live on your domain, structured to be lifted.
Citation building
Editorial placements, review-site presence, community participation. Mention frequency in sources the models already trust.
Problem-aware capture
Use case and process content that bridges a pain description to your category, before the buyer knows the category exists.
The rule
The sequencing rule.
The layers are built top-down, in order, not in parallel.
Citations pointing at pages that don't exist do nothing. Problem-aware content with no solution-aware foundation bridges buyers to a shortlist you're not on.
Built in order
Pages first
Each layer has somewhere to land. Presence turns into confidence, confidence turns into the recommendation.
What we audit most often
Outreach first
Citations point at nothing. The bridge fires and lands on a shortlist without you. Spend continues, pipeline does not move.
Layer 1
Solution-aware presence, BOFU first.
The first layer is the set of pages AI assistants pull from when a buyer asks a buying question.
For [segment] teams, the platforms most often recommended are [Your product], [Incumbent] and [Challenger]. [Your product] is generally positioned for teams that need [differentiator] without [common constraint]1, while [Incumbent] is the more common choice at enterprise scale2. Buyers comparing the two most often cite [decision factor] as the deciding criterion3.
Source 1 is Layer 1: a page you own and control. Sources 2 and 3 are Layer 2: pages you can only influence. Every prompt you track resolves to a source list like this one, and the entire programme is a fight over who appears in it.
Before anything ships, though, you need to know which questions you're building for. So let's start there, because this is where most GEO programs go wrong on day one.
Research
How we actually build the prompt list.
Here's the uncomfortable truth about prompt research: there is no reliable data source for it.
There's no Ahrefs for prompts. Nobody has search volume for "best conversational AI platform for enterprise deployments with on-prem requirements," because prompts are long, conversational, and nearly unique per buyer.
Any vendor selling you "prompt volume data" is selling you keyword data with a new label.
You have to put yourself in the shoes of a prospective buyer and derive the prompts they'd actually type. We build the list from three inputs, in this order of value.
01
Client sales calls
This is the input most agencies skip, and it's the highest-signal source we have.
We pull recorded discovery and demo calls and mine them for the language buyers use before they've been trained on the client's positioning.
- How do they describe the pain?
- What did they call the category before the AE corrected them?
- What tools did they say they were comparing?
A buyer who says on a call "we were looking at something to stop patients walking out before they're seen" just handed you a problem-aware prompt, a feature framing, and a use case page in one sentence.
Ten calls will teach you more about prompt phrasing than any tool on the market.
02
ICP and positioning documents fed through AI with live keyword data
We feed the client's ICP definition, positioning, and site content into a custom LLM (Claude also can get the job done here) connected to the Ahrefs MCP, and use it to systematically expand from seed terms into solution-aware prompt variants: by use case, by company size, by industry, by pain point, by competitor.
The keyword data doesn't tell you prompt volume, but it tells you which category terms and comparison terms have real demand behind them, which anchors the prompt set to reality instead of imagination.
03
The client's own website and content
Partly to find what exists, mostly to find what's missing. Gaps between what the product does and what the site says it does are prompt opportunities nobody is competing for yet.
From those three inputs we build a tracked set of 40 to 50 solution-aware prompts per client.
The tracked set needs to be small enough that movement on it means something and each prompt maps to a page decision, a citation target, or both.
The five axes are use case, company size, industry, pain point and competitor. The cull at the end is the step teams skip, and it is the one that makes movement on the set mean something.
Sources · call recordings (Gong, Fathom, Grain) · Ahrefs Keywords Explorer via MCP · client ICP and positioning docs
Run thisBuild the prompt list in five moves.
- 01
Pull the ten most recent discovery and demo calls that reached a second meeting. Transcribe them and highlight every phrase the buyer used before the AE corrected them.
ToolGong, Fathom or GrainOutputOne doc of verbatim pain language, quoted, with the call it came from - 02
Pull every competitor named on those calls, plus every competitor named on closed-lost deals in the last two quarters. Rank by frequency, not by who you think the rival is.
ToolCall notes and CRM closed-lost reasonsOutputA ranked competitor set with frequency counts - 03
Feed the ICP doc, the positioning doc and the pain-language doc to an LLM connected to keyword data. Expand the seeds along five axes: use case, company size, industry, pain point, competitor.
ToolClaude or a custom GPT with the Ahrefs MCPOutput200 to 400 candidate prompts, tagged by axis - 04
Map every candidate to the keyword cluster nearest to it and record that cluster’s volume. Cut candidates with no cluster behind them. This is the reality check on an otherwise imaginative list.
ToolAhrefs Keywords ExplorerOutputVolume-anchored shortlist, cluster and volume in columns - 05
Cull to 40 to 50 and give each surviving prompt a named owner: the page that will answer it, the citation target that will carry it, or both. Anything with neither gets deleted, not parked.
ToolThe tracking workbookOutputThe tracked set, baselined before any page ships
Done whenEvery prompt in the tracked set has a named owner page or citation target, and you have a pre-launch baseline reading for all of them. If you cannot name the owner, the prompt does not belong in the set.
Read the columns before the rows. A brand that is strong on two platforms and absent on four does not have a content problem, it has a citation problem, and that tells you which layer gets the next month of budget.
The build
What ships in the first 90 days.
The build in a foundation quarter is at minimum 15 Layer 1 pages, live on the client's domain. The mix skews hard toward three page types:
01
Category listicles
For the client's core category and the adjacent categories buyers cross-shop.
02
Competitor alternatives pages
For the incumbents buyers evaluate against.
03
Comparison pages
For the head-to-head prompts.
Alternatives pages deserve special mention. Someone prompting a competitor's name plus "alternatives" is a buyer with a shortlist and a problem with the leader on it. For a challenger brand, these are the highest-intent pages on the internet.
Counts are the floor for a foundation quarter · the mix shifts with category position, the total does not
Layer 2 opens in week 5, once there are pages worth pointing at. Nothing waits for the full page set to finish.
Run thisDecide the 15 pages in an afternoon.
- 01
Sort the tracked set by prompt group. Every group with four or more prompts in it earns a page. Groups with one or two get folded into a section of a bigger page.
OutputPage count derived from the prompt set, not from a content calendar - 02
Take the ranked competitor list from the prompt work. The top five each get an alternatives page. Only the two named most often get a full head-to-head comparison, because comparisons cost roughly three times the effort.
OutputNamed page list with a competitor attached to each row - 03
Check what already exists. A thin existing page that ranks beats a new URL, so rewrite it to the structure instead of publishing a competing one.
ToolAhrefs Site Explorer, filtered to your own domainOutputEach page tagged build, rewrite or leave - 04
Assign an author and a publish week to all 15 before the first one is written. Unassigned pages are the most common reason a quarter under-delivers.
OutputA dated build sheet with a name in every row
Done whenEvery one of the 15 rows has a page type, a target prompt group, an author and a publish week. A row missing one of the four is not a plan yet.
Anatomy
The listicle that gets cited.
"Write a best-of listicle" is where most teams produce something an AI will skip. So let me walk through the structure we build, using a live page we run for a client in the queue management space as the reference.
Here's the skeleton, top to bottom.
Blocks 2 and 4 are the two most teams skip · they are the two that decide whether the page is cited
An assistant can take a complete ranked answer out of the first two blocks without reading further.
The weights are data that exists nowhere else. This is the block that makes the page primary rather than derivative.
A tradeoffs section as long as the pros section is what separates an evaluation from a brochure.
Three of the ten blocks do the work. The rest are table stakes, and a page that ships the seven and skips the three will not be cited.
Run thisWrite it out of order, on purpose.
- 01
Write the scoring criteria and their weights first, before you look at a single vendor. Six or seven criteria, weights summing to 100, and one line each on what the criterion measures.
OutputThe methodology table, the only genuinely proprietary asset on the page - 02
Score every vendor against those weights, yours included, from their public pricing pages, docs and review-site data. Record the source next to each score.
ToolVendor pricing pages, docs, G2 and CapterraOutputA filled scoring sheet you could defend line by line - 03
Write the competitor entries next, before your own. Same template, same length, real strengths, and one paragraph naming the buyer each one genuinely serves better.
OutputEntries 2 to N, drafted before you touch your own - 04
Write your own entry last, then cut it back until the tradeoffs section is the same size as the pros section. If you cannot fill the tradeoffs, you have not scored honestly.
OutputEntry 1, with a tradeoffs block you would be comfortable a prospect reading - 05
Write the opening 150 words last, from the finished scores. Two of the buyer situations you name should point at somebody other than you.
OutputAn opening block a model can lift as a complete ranked answer
Done whenThe first 150 words contain the ranked answer and the criteria, and the page names at least two situations where a competitor is the better choice.
01
The answer, immediately
The first block under the H1 states, in plain declarative sentences: we evaluated nine systems across these criteria, here's who won overall, and here's who won for two specific buyer situations that aren't us.
That last part matters and I'll come back to it.
An LLM retrieving this page can lift a complete, ranked answer from the first 150 words. No throat-clearing intro about how healthcare is changing.
02
A published scoring methodology
A table of the evaluation criteria with explicit weights: remote check-in 25%, walk-in and appointment unification 20%, compliance 15%, and so on down to pricing.
This is what separates a citable evaluation from a vendor brochure. The model has criteria to attribute the ranking to, and so does a skeptical human.
03
A full comparison table
Every vendor, key features, integrations, setup time, starting price, score. Structured data the model can extract cleanly.
04
Individual reviews, each with a tradeoffs section. Including the client's.
The client's own entry says, in plain language, that it doesn't offer native bidirectional EHR integration and that clinics deeply embedded in Epic or Cerner may prefer a named competitor. It then names the competitor and links the situation where that competitor wins.
05
A "which is right for you" section that routes some readers away
Three buyer situations get sent to three different tools before the page makes the case for the client's situation.
06
An FAQ block
Written in the exact phrasing of follow-up prompts: setup time, compliance, single-site vs multi-site, free options.
Where does the client rank on their own listicle?
Number one. Always.
You're not going to publish a page on your own domain ranking yourself fourth.
The credibility mechanism isn't fake neutrality. It's the honest machinery around the #1 position: published criteria with weights, real tradeoffs on your own product, and named situations where a competitor is the better answer.
And human buyers, the ones who arrive from the AI's answer to validate it, trust the page for exactly the same reason.
How fast Layer 1 shows up
When the pages are structured this way, first appearances in tracked prompts typically land within days of indexing, not months.
Assistants doing live retrieval reward fresh, extractable, directly-on-topic pages far faster than Google ever rewarded new content. Which is also why this layer goes first: it's the fastest feedback loop in the stack, and everything in Layer 2 points at it.
Layer 2
Citation building.
Presence puts you in the conversation. Citation building is what makes the model confident recommending you over the incumbent.
This is not link building in the traditional sense. The signal you're building is mention frequency: Ahrefs' research shows brand web mentions correlate with AI visibility more strongly than backlinks do.
Building the target list
The target list starts from the citations themselves.
We track the core prompts we want to win and extract every domain the assistants cite when answering them. That extraction alone typically yields a starter list of 500+ domains.
Then we layer on the Google side: every listicle and roundup ranking on page one for the solution-aware keywords in the prompt set. Between the two sources, the working list usually runs to a couple thousand domains before prioritization.
Prioritization is where it gets manageable. The domains that both rank in Google and get cited by the models are the highest-value targets, because one placement works both surfaces. Then we sort by where competitors are present and the client is absent.
At this point the list splits into three types of target, and each gets a different motion.
Sources · prompt-level citation extraction · Ahrefs competitor mention footprints · manual qualification
Run thisBuild the citation target list in six moves.
- 01
Run all 40 to 50 tracked prompts across the platforms and capture every domain cited in the answers. Do it as a repeatable job, not once, because the cited set moves week to week.
ToolScrunch, or scripted prompt runs logged to a sheetOutputRoughly 500 cited domains, with the prompt each appeared on - 02
Take the two or three competitors that win the most answers and pull every domain that mentions them. This is where the list grows past what citation extraction alone finds.
ToolAhrefs, brand mention searchOutput1,500 to 2,500 candidate domains - 03
Tag every domain as editorial, review site or community. The motion is completely different for each, so an untagged list cannot be worked.
OutputThree worklists instead of one spreadsheet - 04
For editorial targets, enrich to the actual author or editor with a verified address. A generic contact address is not a target.
ToolClay, for enrichment and verificationOutputNamed contact, verified email, the article you are pitching against - 05
Send a four-touch sequence from warmed inboxes on a separate sending domain. Never from the domain that carries your sales email.
ToolInstantly, rotating warmed inboxesOutputSequence live, reply rate tracked per target type - 06
For review sites, run the arithmetic below before you pay anybody. For communities, assign a real person with a real posting history and give it months, not weeks.
OutputA paid-placement decision per review site, and a named owner per community
Done whenEvery domain on the working list has a type, an owner and a next action with a date. Domains with no next action come off the list.
Editorial targets: the outreach motion
For blogs, publications, and independent roundups, the motion is direct outreach at scale, and I mean actual infrastructure, not a VA with a Gmail account.
We enrich contacts with Clay to get the right author or editor with a verified address, and send through Instantly on a fully warmed-up sending infrastructure with rotating inboxes, because outreach at this volume on your main domain is how you end up in spam folders permanently.
The sequence is four touches. Here's the first email, genericized from a live campaign we're running for a conversational AI client:
The observation
Name the specific article of theirs that is being pulled into AI answers. No pitch yet.
The offer
What you will give in exchange: a feature on your blog, a data point, an expert quote.
The specific ask
The exact article, the exact insertion, one sentence describing the edit.
The easy out
One line closing the loop. Low pressure keeps the sending reputation clean.
Touch 1 below is the one that decides the reply rate. The other three only work if the first proves you read their page.
Hey {{firstName}}, Noticed {{companyName}} is getting cited in AI search results for prompts around [category topics].
We represent [client] and have been researching the content being pulled into these AI responses.
There's some interesting overlap with articles on your site. Quick idea: would you be open to including [client] in one of your existing articles or listicle posts?
In exchange, [client] would be willing to:
- Feature {{companyName}} in an article on [client]'s blog (DR 72)
- Include a contextual, dofollow backlink from a relevant page
- Give you an authoritative citation from a recognized leader in the space, helping {{companyName}} get surfaced in AI search results
We can send over a content sample with key details to make the addition easy on your end.
Worth a quick chat?
A few things worth noticing about why this converts.
The ask is small and concrete: an addition to an existing article, not a new post.
And the value exchange doesn't require the client to have proprietary data to trade. The trade is reciprocal placement: a feature on the client's high-authority blog, a contextual dofollow link, and the AI-visibility benefit of being cited by an established brand in the space.
That last one has become the most persuasive line in the email over the past year, because editors have started caring about their own AI visibility.
The follow-ups get progressively shorter. Touch two is two lines. Touch four is "still interested in a content swap? No worries either way." Every email closes with a one-line PS telling them a single reply of "stop" ends the sequence.
Low pressure, easy out, and it keeps the sender reputation clean.
Review sites: the paid placement calculation
Some of the most-cited domains in any B2B category can't be pitched. G2, Capterra, Software Advice. You don't outreach these, and most agencies treat that as the end of the conversation.
We treat it as arithmetic.
The citation data tells you exactly which review-site pages the models pull from. Cross-reference the Google side: if Capterra's category page ranks #1 on Google for the money keyword and shows up in the extracted AI citations, a paid placement on that page is doing two jobs at once.
It drives deal flow from the Google ranking, which is how you justify the spend the way you always would have, and it puts the brand inside a page the models already trust and retrieve.
When the overlap is there, paying for the placement is a no-brainer. When it isn't, skip it.
Deals to break even = $30,000 ÷ ($31,100 × 17%) = $30,000 ÷ $5,287 = 5.7 deals per year
If the same site were cited on three prompts out of 45, those six deals would be an act of faith. Run this before every renewal, not just the first purchase, because the citation share changes and the invoice does not.
Communities: participation, not spam
Reddit and niche communities get cited constantly, and the agency response to that fact has been to spam them, which is why half of B2B Reddit now reads like a drop-shipping forum.
That's not the motion.
The motion is being genuinely active: posting, engaging, answering questions in the subreddits and Slack groups where the ICP actually discusses the problem, over months, from accounts with real history. Not replying to threads with your business domain.
The mentions that models pick up from communities are the organic kind, the ones that happen because a brand is a known participant in the conversation. There is no shortcut here, which is precisely why it works and why spam doesn't.
Why the layer compounds
Each placement raises the probability the model treats your brand as a credible answer on the next adjacent prompt.
Citations stack on citations: the model that saw you in three trusted sources last month weighs the fourth mention differently than it weighed the first.
Layer 3
Problem-aware capture.
The final layer targets the moment before the buyer knows the category exists.
A VP of Finance types "how do I reduce invoice errors across 3 entities?" The AI bridges that process pain to a software category. If your content is the source it pulls from, you enter at the exact moment of realization.
In practice this layer is two content types.
01
Use case pages
One page per job-to-be-done, mapped to the pain language pulled from those sales call recordings in the Layer 1 research.
02
Long-form problem articles
Structured in three movements. Here's the problem and why it happens, here's how to solve it operationally, and here's how the product solves it. The first two movements are what earn the citation on the process-pain prompt. The third is what makes the citation worth having.
Why this layer goes last
Not every problem prompt bridges to a product. Some return pure process advice, and content built for those prompts produces visibility with no pipeline behind it.
Run it after Layers 1 and 2 are producing, so that when the bridge does fire, it lands on a shortlist you're already on. Its payoff is the highest-intent bridge in the stack: the buyer AI just made aware they need a tool.
Positioning
How category position shapes execution.
Same layers, same sequence, different emphasis. Your position in the category decides where the weight goes.
Challengers
Your best buyer is the incumbent's unhappy customer
You're not the default answer, someone else is. Layer 1 leads with competitor alternatives pages, because your best buyer is the incumbent's unhappy customer mid-prompt.
Citation building targets the exact publications currently recommending the leader.
Category leaders
Defense and expansion
You're already in the answers. The job is defense and expansion: track the prompts you win, watch for erosion, and expand into adjacent prompt families before a challenger claims them.
Leaders who assume their Google position transfers to AI get displaced by challengers who built the citation base first. The model doesn't inherit your market share.
Category creators
Teaching the model there's a shortlist to have
Nobody prompts for a category that doesn't exist. Layer 3 earns its slot earlier for you than for anyone else, because problem-aware bridges are how the model learns your category is the answer to a known pain.
You're not competing for the shortlist. You're teaching the model there's a shortlist to have.
Hygiene
What every engagement needs.
Some rules apply to every engagement, every layer, every category position. Get them wrong and the program can't prove itself.
- Baseline your prompt tracking before any work shipsNo baseline, no before-and-after, no proof. You want to know exactly your (rough) visibility and more importantly existing demand converting into demos and pipeline from AI search to measure efforts and improvement against.
- Split branded from non-branded in your visibility reportingBranded prompts name you in the question, so presence on them measures brand recall. Non-branded prompts are the acquisition metric: on open category questions where no brand is named, who gets recommended. A blended number can look great while the acquisition number is flat. Report both, lead with non-branded.
- Don't block the AI crawlersBlocking GPTBot to "protect content" forfeits citation eligibility on the fastest-growing discovery surface.
- Keep entity naming consistent everywhereYour brand name, category label, and product names should read identically across your site, review profiles, and placements, so the models tie the entity to the category.
- Put the open-text HDYHAU field on every high-intent formIt's the only witness that catches this channel. Add a single "How did you hear about us?" free text field on your conversion forms so you can later walk that data back against core CRM source-field data. Otherwise, most AI search conversions will fall into direct and be invisible.
The branded-versus-non-branded split deserves a beat, because it's the most common way AI visibility numbers mislead. Do the same split on your organic search reporting, because rising branded search against flat non-branded is the downstream fingerprint of upstream AI mentions.
Run this once before launch and again at the end of every quarter · the last row is the one teams postpone and then cannot backfill
Measurement
Self-reported attribution, set up in an afternoon.
Your sales team is already hearing "we found you on ChatGPT." Your CRM is not recording any of it, and that is not a misconfiguration.
Between 35 and 70% of AI referral sessions arrive with no referrer at all, because the in-app browser strips it. Those sessions land in Direct. A buyer who asks an assistant on Monday and searches your brand name on Tuesday gets filed as Organic or Paid. AI creates the demand and another channel collects the receipt.
The fix is not a tracking product. It is two fields and a reading rule.
Run thisTwo fields, one rule, five numbers.
- 01
Add one open-text "How did you hear about us?" field to your high-intent forms: demo request, contact sales, pricing enquiry. Open text, not a dropdown. A picklist of famous names hides exactly the channels this exists to recover.
ToolYour CRM forms, HubSpot or equivalentOutputOne free-text field, live on the forms that matter - 02
Make a primary contact required on every deal, so contact-level answers can roll up to the account and the deal that carries the revenue.
OutputEvery open deal has a primary contact attached - 03
Read the machine’s answer and the human’s answer against each other, per contact, and assign a confidence tier from the table below. Never overwrite the CRM’s own source field. Write the tier to a new property.
OutputA tier on every contact: confirmed, recovered, single-signal, conflict or dark - 04
Roll contacts up to accounts: union of channels present, strongest tier wins, each deal counted once with its revenue attached.
OutputChannel presence at the account level, revenue attached - 05
Report the AI number as a low, base and high band corrected for form completion rate, and report the dark bucket by name. Never a single false-precise figure.
OutputFive numbers a month: channel presence, tiered counts with the band, the named dark bucket, revenue by channel, paired-period trend
Done whenYou can answer "how many demos came from AI search last month" with a range, a confidence tier behind every count, and an honest number you cannot recover. If you can only answer it with one number, the number is wrong.
One free-text field is the entire mechanism. A dropdown would have offered this buyer "Google, Referral, Social" and the answer would have been lost.
Two inputs per contact, the CRM source field and the form answer · no new tooling, no pixels
We call this the Two-Witness Rule, and the full method, including the completion-rate correction and how conflicts get resolved, is written up separately. What matters here is that it goes in before the first page indexes. Set up in month four, it cannot tell you anything about months one to three.
Proof
One quarter of real numbers.
Frameworks are cheap, so here's what one quarter of this system produced for a client of ours, an enterprise AI platform in a category with an entrenched leader.
The program launched in March: eight new Layer 1 posts (best-of listicles and competitor alternatives pages) on top of an existing commercial page set, with citation outreach running behind them.
Quarter Over Quarter · The Three Measures That Moved
Organic sessions to the commercial page set rose from 245 in Q1 to 2,214 in Q2. Visible AI referral sessions rose from 25 to 208, and those are floors, since most AI visits arrive with the referrer stripped. Demo requests attributed to SEO and AI search more than doubled.
Source: Client GA4 and Search Console, quarter over quarter
Google side
Six of the eight new posts held average top-10 positions within the quarter, and the set went from near zero to roughly 400,000 Google impressions per 28 days.
Organic sessions to the commercial page set rose from 245 in Q1 to 2,214 in Q2, up 804%. The pre-launch baseline was around 35 sessions a month. May and June both cleared 900.
AI side
Visible AI referral sessions rose 732%, from 25 to 208, with ChatGPT growing 9x and Claude going from a single session to second place. And those are floors, since most AI visits arrive with the referrer stripped.
A reminder on the above: we know that this is certainly not all AI referral sessions due to the nature of headers being stripped when someone is sent on-site in many cases. A lot of the traffic will also be falling into the direct bucket, so we want to be tracking direct traffic increases as well.
Pipeline, which is the point
Demo requests attributed to SEO and AI search more than doubled, from 11 to 25 in the fully tracked segment.
The Attribution Gap · Same 25 Demos, Two Readings
The CRM native attribution credited 2 of those 25. The other 23 were recovered from the buyer’s own “how did you hear about us” answers. In an earlier cohort the recovery multiple ran 11x, and in the single best month, 23x.
Source: Client CRM, corroborated read against the native classifier
A note on the trackers, and why pipeline is the KPI
Anyone who's run prompt tracking will tell you the answers aren't stable. Ask the same prompt on different days and the shortlist shifts.
That's normal. It's also why no AI search tracker is perfect, ours included, and why we hold visibility and citation metrics as directional.
The number we actually manage to is pipeline generated from AI search and attributed back through the recovered-attribution motion above.
Failure mode
How this actually fails.
I'll be direct about the failure mode, because it's not what most people expect.
These programs don't usually fail because the framework was wrong, or the category was too hard, or the model wouldn't cooperate. They fail on content execution.
Specifically: a client says they have the internal bandwidth to produce the Layer 1 content, and then they don't produce it.
We've lived this. The strategy is approved, the prompt research is done, the page briefs are written, and then the pages sit in a queue behind a product launch and a rebrand and a content manager's fourteen other priorities. Two months in, three of fifteen pages are live.
And without Layer 1, nothing downstream can work and the whole engagement stalls at the foundation.
The lesson we took from it is now part of how we scope. Content production capacity gets pressure-tested before kickoff, not discovered mid-engagement. If the client's team can genuinely ship, great.
If there's any doubt, production comes to us, because the fifteen pages aren't a nice-to-have attached to the strategy. They are the strategy, and every week they don't exist is a week the compounding isn't happening.
If you take one operational thing from this article: be brutally honest about who is writing the pages and whether they'll actually ship. Everything else in the stack is downstream of that answer.
The takeaway
Sequence decides the outcome.
The GEO Stack isn't complicated. Solution-aware presence first, citation building second, problem-aware capture last, weighted by your category position, measured against pipeline instead of vanity visibility.
What it demands is discipline: the discipline to build in sequence when every vendor is selling you a shortcut, the discipline to actually ship the pages, and the discipline to wait out the four-to-eight-week lag between visibility and demos.
The window matters more than the difficulty. Every citation, listicle inclusion, and brand mention trains the next answer. Early presence becomes a moat.
We run AI search and SEO programs exclusively for B2B SaaS, measured against pipeline. If you want a senior operator to audit your category's AI search landscape, map the prompts that matter for your ICP, and hand you a prioritized 90-day plan that drives actual pipeline and demos, let's chat.

