Most SEO forecasts are one made-up number multiplied by another made-up number. Somebody pulls total search volume out of Ahrefs, multiplies it by a conversion rate that looked about right, and calls the result a projection.
There's no ramp in it, so the math quietly assumes you rank in month one. Nothing in it accounts for how deals actually close, so it never gets as far as revenue. And it doesn't answer the question your CFO is going to put to you in the budget meeting, which is what this turns into in dollars and in which month.
I'd rather walk through ours. Not the philosophy of forecasting, but the five steps in the order we run them and the assumptions we won't fudge.
We've built this now for process automation platforms, wallet infrastructure, banking analytics, franchise ops software, time tracking, and ATS products. The categories change every time and the sequence doesn't.
Two things separate it from most forecasts you have been handed. You approve the target list before we model anything against it, and once we commit, the only scenario we hold ourselves to is the conservative one.
The first three steps happen before anyone signs anything. The last two are what we do once the work is running, and they exist to make sure the number we committed to in step 3 is the number we actually get measured on.
Step 1
Map the prompts.
Everything downstream depends on the denominator, so if the target list is wrong then the rest of the model is just precision applied to the wrong problem.
We go find every solution-aware term a buyer in that category actually uses, on both sides of the split.
The Google side
Ahrefs, filtered hard
Filtered to the target country, branded terms stripped out, irrelevant parent topics cut. We're looking specifically for the phrasing people use once they already know the category exists and are choosing between vendors, so that's best, top, alternatives, versus, pricing, comparison, and competitor names.
The AI side
Sales call transcripts before tools
The prompts buyers type into ChatGPT, Perplexity, Gemini, and Claude at that same moment in the journey. Those come out of sales call transcripts before they come out of any tool.
Mining your own recorded calls for how buyers describe the problem in their own words is the highest-signal input in this entire process, and almost nobody bothers to do it.
Both lists then get clustered across category queries, use cases, industries, company sizes, pain points, features, and comparisons.
The finished map lands somewhere between 50 and 100 terms, and that ceiling is intentional. You can generate a thousand keywords in ten minutes, but a forecast built on a thousand terms is a forecast you have no intention of executing against.
We also cut any prompt that doesn't bridge to a product recommendation. A lot of problem-aware queries return pure process advice with no vendors named and no shortlist offered, and while those are real prompts that people really type, they have no business anywhere near a revenue model.
Row four carries a visible estimate tag because no public volume data exists for AI prompts. Row five is real volume that converts at a fraction of the rows above it, which is why every row is weighted by stage before it touches revenue.
Do this todayBuild a target list you can defend to your CEO
- 01
Pull the Google side from Ahrefs, one country, branded terms stripped, irrelevant intent removed.
ToolAhrefsOutputRaw term list - 02
Build the prompt side by hand from how buyers actually ask, then cut any prompt that never bridges to a product recommendation.
ToolManual · call notes - 03
Cluster both lists by category, use case, industry, size, pain point, feature, comparison. Stop at 50 to 100 terms.
ToolClustered map - 04
Tag every row by funnel stage. Bottom-funnel gets weighted materially higher before any revenue math runs.
ToolStage column - 05
Have your own sales team sign off on the list in writing before anyone builds a formula on it.
ToolApproved list
Done whenSomebody who sells into that market every day has read the list and said yes, that is what our buyers ask.
Step 2
You approve the list before we model anything.
This is the step almost nobody runs, and it's the one that makes everything after it defensible.
Before we build a single formula, the map comes to you for review. You read it, you tell us what is missing, you tell us what is off, and you confirm these are the terms your buyers actually use.
Usually we hit it close to exactly, which is the whole reason the step exists. When the person who sells into that market every day reads a hundred prompts and says yes, that is the list, our research has become your assertion.
Every argument that could happen later in the engagement traces back to this moment. If your CEO asks in month five whether we are chasing the right terms, the answer is a document you approved before any money changed hands. And when you want to add something, which happens often enough, it goes straight in and the model absorbs it.
This note is the cheapest insurance in the engagement. If your CEO asks in month five whether the right terms are being chased, the answer is a document you approved before any money moved.
Step 3
Run it through the model.
Now it turns into math. We take total addressable volume across the approved map, ramp it across twelve months, run it at three scenarios, and benchmark the whole thing against a 3x ROI target.
The inputs come from you
The conversion economics come from your own systems, and we pull them live off your GA4 and CRM while we are on the call.
- Sitewide visit-to-demo or visit-to-trial ratePulled from your own GA4, not a category benchmark.
- Demo-to-customer or trial-to-paid close rateStraight out of your CRM.
- ACV or LTV, depending on how the contracts are structuredWhichever your revenue model actually runs on.
- The monthly investment and their current organic baselineSo the ROI line is measured against where you are, not zero.
We use your numbers even when they look low, because a model built on your own data survives a finance review and a model built on industry averages gets torn apart in one.
That principle costs us something on almost every build. One banking analytics client had a high ACV on multi-year contracts, and we modeled year one revenue only, which understates the actual deal by roughly two thirds. It would have been easy to run the full contract value and produce a much prettier ROI. We wrote the limitation into the document instead.
A process automation client had three funnel stages where our template assumed two, so we rebuilt the model around their stages rather than flattening their funnel into the shape ours expected.
And when a number genuinely doesn't exist, we don't borrow one and hope nobody notices. One franchise ops platform prices per location and publishes nothing, so we triangulated an ACV from location counts and category pricing, put it in the model labeled as an estimate, showed the arithmetic and the sources in the definitions tab, and flagged it as the single biggest lever in the workbook and the first thing worth pressure-testing.
Every cell here is editable in the file you receive. Dragging the capture rate around for ten minutes yourself is the moment the model stops being our pitch and starts being your plan.
Bottom-funnel terms carry more weight than middle-funnel
Every term on the map gets tagged by funnel stage, and bottom-funnel terms are weighted to convert materially better than middle-funnel ones. Somebody searching "[competitor] alternatives" has a shortlist and a problem with the leader on it, and treating that visitor the same as somebody reading a definition throws the model off before it starts.
That weighting is a large part of why two agencies can forecast identical traffic and land multiples apart on revenue. Traffic on its own isn't the unit worth forecasting against. Weighted traffic is.
AI prompt volume is an estimate and we say so
There's no public volume data for AI prompts, so anybody quoting you one has invented it. We assign each prompt a share of the Google cluster volume it maps against, and that share sits in a labeled, editable cell rather than buried three layers into a formula where nobody can argue with it. Clients with unusually heavy AI referral share get a higher one, and we tell them why.
The timing problem
This is where most forecasts fall apart, and it took us several rebuilds to get right, because there are two separate timing problems and most models only solve one of them.
Problem one
When a term starts ranking at all
Terms enter the model staggered by priority rather than all switching on at once, and the AI prompts trail their Google equivalents, since citation velocity lags publication.
Problem two
What a term does after it starts ranking
This is the one people miss. An early version of our model had keywords jumping from zero to full volume in the month they turned on, which produced stair-step plateaus and a growth line that sat flat for months at a stretch.
Real keywords mature gradually, so every term now climbs its own maturity curve across its first year of ranking, starting at a small fraction of available volume and reaching full share near the end.
Stack both curves together and the projection starts to look like SEO rather than something drawn freehand.
$14,400Month twelve, from one term, at realistic capture. Fractional customers are never recognised, which is why the curve steps rather than glides.
Two agencies can forecast identical traffic and land multiples apart on revenue. The stage weight and the whole-customer rule are where most of that gap lives.
Three scenarios, one lever
Conservative, realistic, and optimistic differ on traffic capture rate, meaning the share of addressable volume the program actually takes at full ramp. Nothing else moves between them. Once scenarios start differing on six inputs at once, nobody can reason about the gap between them.
The capture rates themselves shift with the shape of the opportunity. A small map against a large ACV gets more optimistic capture assumptions than a large map against a small one, because taking a quarter of sixty terms is a different proposition from taking a quarter of three hundred.
Everything you could reasonably question is an editable input cell, including capture rate, conversion rate, close rate, ACV, investment, and the maturity curve. Change any one of them and the full twelve months recalculates in front of you.
Nothing except capture rate moves between the three. Once conversion, close rate and ACV are your own numbers, changing them per scenario would be inventing a different company three times.
Step 4
Hold the line at conservative.
Everything up to this point happens before a contract exists. Now the work starts, and the forecast stops being a sales document and becomes the thing we get graded against.
We publish all three scenarios and commit to one.
Three Scenarios, One Commitment
The three scenarios differ on one input, the share of addressable volume the program captures at full ramp. Conservative is the only line we agree to be measured against; the other two exist so the client can see the shape of the upside.
Source: Client GA4 and HubSpot, 12-month model output
Conservative is the number that goes in the proposal, becomes the 90-day baseline, and would embarrass me to miss. Realistic and optimistic are there so you can see the shape of the upside, not so we have something to point at in month six.
The metric we commit to is demos and pipeline, not visibility and not rankings.
We do report leading indicators, because in the first few weeks they're the only honest thing there is to report:
- AI search visibility and movementAcross the core solution-aware prompts.
- Citation rateHow often your own content shows up as a clickable source in AI answers.
- Google rank on the core solution-aware keywordsThe Google-side half of the same map.
- Traffic to the specific new bottom-funnel URLsWhich are net new, so anything they produce is program-attributable by definition.
Those are early-warning signals rather than the report card. Somewhere between week four and week eight the conversation moves to demos and pipeline, partly because that's what we're being paid for and partly because that's when the first bottom-funnel pages start producing.
Four to eight weeks also happens to be the lag we consistently observe between a citation or a ranking landing and a demo showing up on the other end. So the leading indicators aren't filler, they're what tells you how the next two months of pipeline are going to look. They just never get to be the deliverable.
Step 5
Layer the attribution on top.
Which raises an obvious problem, and it's the one that sinks most AI search engagements.
We have just committed to a demo number. Your CRM cannot count those demos.
HubSpot builds its source field from the referring URL. When ChatGPT, Perplexity, Gemini, or Claude sends a buyer over, the referrer gets stripped, so the visit arrives with no source data attached and lands in the system as Direct.
On one 90-day program, 13 buyers typed into the form that they'd found the client through AI search, and HubSpot machine-tagged 2 of them.
What The CRM Could See
Commit to a demo number and let this dashboard referee it, and you have agreed to be judged by an instrument that structurally cannot see most of your work. This is why the attribution goes in during week one, before there is anything to measure.
Source: HubSpot contact records, corroborated read
Sit with what that means for a second. If we agreed to a demo target and then let your dashboard be the referee, we'd have been held to a count that was structurally missing most of the demos we generated. So would any agency in our position, and plenty of good programs have been cut over exactly that.
There's a second-order version that's sneakier still. Plenty of buyers see you recommended in an AI answer and never click at all. They remember the name and come back a week later through branded search or by typing the URL, which lands in your dashboard as branded organic or Direct traffic even though it's the downstream signature of an upstream AI mention.
So we set this up in week one, before there's anything to measure. Every client runs on a two-witness rule, where a contact only counts as AI-sourced when two independent signals agree, meaning the CRM's own classification and the buyer's self-report in an open text "how did you hear about us" field on the demo form. Where the two witnesses disagree, we report the gap rather than quietly picking whichever number flatters us.
The operating rule we give every client is to treat the dashboard as a floor, treat the form self-reports as the recovered signal, and build the reporting around the distance between the two.
This is the piece that makes step 4 enforceable. Commit to a demo number without it and you've agreed to be judged by an instrument that can't see your work.
Failure modes
The four ways a forecast lies.
Every one of these was a bug in our own model before it became a principle.
Lie 01
Fractional customers
An early build recognized revenue in month one off a fraction of a customer, and you can't close part of a deal. The fix was to accumulate pipeline month over month and only recognize customers on whole numbers, which means conservative now shows zero revenue for the first stretch of the year and first revenue landing mid-year.
The chart looks worse and it's also correct, because a forecast that pays out in month one against a 90-day sales cycle has clearly never met the sales team.
Lie 02
One term carrying the whole model
In an ATS forecast, a single head term accounted for roughly two thirds of the entire keyword universe. It was scheduled to rank late in the year and it spiked the whole projection when it did. We cut it and redistributed the remaining terms across the ramp months.
If pulling one keyword out changes your forecast materially, what you have isn't a forecast, it's a bet on one keyword.
Lie 03
Inflating capture to hit the target
One franchise ops model came back well below the 3x benchmark in conservative, because total addressable volume in that category was genuinely thin. The easy fix would have been nudging the capture rate up until the problem disappeared. We shipped it short of target instead and wrote that the honest lever was confirming a higher real ACV rather than assuming more of the market.
A small TAM is a finding, and findings belong in the deliverable.
Lie 04
Forecasting the wrong output
Visibility scores aren't forecastable and aren't what anybody is buying, so the output columns are demos, customers, revenue, and cumulative ROI.
A forecast whose terminal metric is a visibility percentage can never be held against a CRM later on, which is a large part of why some vendors prefer building them that way.
The real test
When the honest model misses.
Sometimes conservative doesn't clear 3x, and what you do in that moment is the real test.
For one time tracking client, even the optimistic case came in short. There were five levers available, being search volume, site conversion rate, trial-to-paid rate, ACV, and monthly investment.
Hitting the target on trial-to-paid alone would have meant more than tripling a rate that was already at industry norms, which isn't a lever, it's a fantasy, and writing something like that into a model is how you end up in month eight explaining why the plan never had a chance.
The honest combination was expanding a keyword universe that was genuinely under-built, plus a realistic conversion lift on the bottom-funnel landing pages, plus pushing annual plans to raise ACV. That got us close enough to target on assumptions we could defend line by line.
Forecast vs actual
What it looks like when the 90 days are up.
Enough process. Here's one of these models sitting next to what actually happened.
Vertical B2B SaaS, mid-market and enterprise buyers, a 60 to 120 day sales cycle, three entrenched incumbents in the category. Kickoff on the first of March, with an approved map of 89 tracked prompts and their matching Google terms.
The forecast-versus-actual deck we walk through at the 90-day mark has three rows in it. Not fifteen. Three.
Day 90 · Share Of The Realistic Forecast Delivered
Traffic landed at 81% of realistic, a genuine miss that stays in the deck exactly as it is. Demos cleared it by more than six times. Customers came in at six against a model that does not recognise its first whole customer until the second half of the year.
Source: Client CRM at day 90, against the committed conservative scenario
Row 01
Traffic to the new bottom-funnel pages
Conservative called for 319 visits and realistic called for 638. We measured 519, which is 63% ahead of the line we committed to and short of the realistic one.
That's a miss against realistic and it stays in the deck exactly as it is, because the entire reason you publish three scenarios is that you don't get to quietly relabel which one you were aiming at once the results land.
Row 02
Demos
Conservative called for three or four and realistic for a little over nine. Fifty-nine buyers named Google or AI search on the intake form and booked.
Those 59 accounted for 63% of every demo the company booked from any source since kickoff.
Row 03
Customers
Six closed inside the window, against a conservative model that doesn't recognize its first whole customer until well into the second half of the year.
That's the fractional-customer rule from earlier doing its job, and doing it in the direction that favors the client.
Underneath those three rows sat the leading indicators we'd been reporting since week one. Eight bottom-funnel pages live with seven more in production, 23 authority placements, and citation rate across the tracked prompts climbing from 33% to 47%.
Every one of those indicators moved before the demos did, which is the argument for reporting them at all. None of them is what we agreed to be judged on.
Do this todayFive questions to ask any agency's forecast
- 01
Which conversion rates came from our data, and which came from a category benchmark? Point at the cells.
Bad answer"Industry standard" - 02
Where is the ramp? Show me the month the model assumes we start ranking, per term.
Bad answerFlat month one - 03
Which numbers in here are estimates, and are they labelled as estimates in the file?
Bad answer"All of it is data" - 04
Which single scenario are you committing to, and what happens in month five if you miss it?
Bad answerA range, no commitment - 05
How will you count demos when our CRM cannot see AI referrals at all?
Bad answer"HubSpot will show it"
Done whenYou have the committed number, the instrument that measures it, and the list it rests on, all in writing, before any money moves.
The takeaway
Ask for the forecast before the proposal.
The spreadsheet isn't the impressive part, because anyone can build a spreadsheet.
What is harder is an agency writing down the number it expects to be judged against before any money moves, having you approve the target list it rests on, and then measuring itself against that number out of your CRM rather than out of a dashboard it controls.
The gap between nine forecast demos and 59 real ones isn't a brag about the work. It's what happens when you commit to the conservative line and then build attribution honest enough to actually see the channel. A model built this way will systematically understate what it finds, which is the right direction for a model to be wrong in.
If you're evaluating an AI search or SEO program, in-house or agency, ask for the forecast before the proposal. Then ask four things about it.
- Did I approve the target list this is built on?If the answer is no, you are being forecast at rather than forecast for.
- Which cells can I change?A model you cannot poke at is a brochure.
- Where did that conversion rate come from?Your GA4, or an industry average somebody found convenient.
- In which month does this model say I see my first dollar?If the answer is month one, close the file.
We run AI search and SEO programs exclusively for B2B SaaS companies. If you want a senior operator to map your category's real prompt set, put it in front of you for approval, and model it against your own conversion data before you commit to anything, let's talk.

