A prospect forwarded us a proposal for AI search engine optimization services last month, with a case study on page four. Brand coverage up from 41% to 63% in ninety days. Two engines, twelve prompts, a bar chart with no y-axis label.
We asked the vendor one question by email: how many times was each prompt run, and over what window. The answer came back in a day. Once a week, one run, twelve prompts, two engines.
That report was not dishonest. Everything in it happened. It also could not have detected a 22-point improvement, because at twelve prompts the instrument cannot resolve a 22-point difference from nothing at all. The number moved. Nobody could say whether the brand did.
That gap, between work that happened and work that can be shown to have happened, is what this post is about.
Do AI search engine optimization services work?
AI search engine optimization services do work, in the narrow sense that structured answer content, entity consistency and off-domain placement measurably change which sources AI engines retrieve. What usually does not work is the reporting. Most vendor panels are too small to distinguish a real improvement from normal day-to-day variation, so the number moves whether or not the brand does.
The proof ledger
What a reported improvement has to clear before it means anything
Layer one is the precision your panel size can support. Layer two is how much sampling each engine costs you. Layer three is the ceiling your category imposes no matter who you hire.
Layer 1: the 90% band a brand coverage figure sits inside, by panel size
10 prompts tracked
39% to 71% · 32 points wide
A brand whose true coverage is 55% can honestly report anything in this range. Most vendor case studies are built on panels this size.
50 prompts tracked
49% to 62% · 13 points wide
The 10 to 50 step buys 19 points of precision. This is the knee of the curve and the floor worth insisting on.
100 prompts tracked
51% to 60% · 9 points wide
The 50 to 100 step buys 4 more points. Past here you are paying for decimal places, not for decisions.
Layer 2: day-to-day source reuse, 30-day window, seven engines
One cadence cannot serve both ends of that list
Three quarters of what Perplexity cited yesterday it cites again today. A quarter of what ChatGPT cited yesterday survives. A weekly check reports a durable position on one engine and a coin flip on the other, and reports them in the same column.
Layer 3: the ceiling, from our own 90,132-answer corpus
Both halves of that answer matter. Dismissing the category because the case studies are weak is as wrong as buying on the case studies. The mechanism is sound and the instrument is bad, and the instrument is the part you can fix before money changes hands.
A percentage without a panel size is not a result. It is a rumor with a decimal point.
5 reasons the improvement you were shown is probably noise
Before you evaluate any specific vendor, understand what would make almost any report in this category unreadable. These five reasons are structural. They apply to good agencies and bad ones alike.
Reason #1: The panel is too small to resolve the change being claimed
Otterly published the measurement in September 2026 that the category had been missing: 252,407 AI search answers across seven engines, 520 prompts, 40 to 91 days depending on category.
The finding is a precision curve. At 10 tracked prompts, 90% of brand coverage results fell in a band 32 points wide, from 39% to 71%. At 50 prompts the band narrowed to 13 points. At 100 prompts, 9 points.
Read that against the case study on page four. A twelve-prompt panel reporting 41% to 63% is reporting a move that fits comfortably inside its own error bar.
Reason #2: One cadence is applied to engines that behave nothing alike
The same study measured how much of what an engine cited yesterday it cites again today. Perplexity reused 74.9% of its sources day to day. ChatGPT reused 25.8%.
That is close to a threefold difference in how sticky a win is. A weekly check on Perplexity is close to the truth. The same weekly check on ChatGPT samples a surface that has replaced three quarters of its sources since the last reading.
Most reports average the two into one line and call it AI visibility.
Reason #3: The vendor wrote the prompts that define the score
A preprint submitted on September 6, 2026 by Olivier Martinez makes the theoretical version of the argument: prompt corpora define an "answer market" that need not represent actual user demand.
In practice that means the score is a function of the prompt set, and the prompt set was chosen by the party being graded on the score. Ask any vendor to add ten prompts you wrote yourself and watch what happens to the baseline.
Reason #4: Nobody in the category publishes uncertainty intervals
Machine Relations audited 16 AI visibility platforms on methodology disclosure in September 2026, scoring nine dimensions including sampling volume, deduplication rules and uncertainty treatment.
Uncertainty intervals and exact deduplication rules were the two fields most often returned as undisclosed. Those are precisely the fields that decide whether a reported change is real.
The audit is vendor-run and omits its own author, so take the dimension list and leave the rankings. The dimension list is the useful part.
Reason #5: Mentions, citations and recommendations get blended into one number
Being named in prose, being cited as a linked source, and being recommended as the pick are three different commercial events with three different values. Blend them and a vendor can raise the headline by getting you mentioned more often in answers that still recommend a competitor.
This is the single easiest number to move without moving anything a buyer cares about. It is also the one most likely to be reported without a definition.
What a vendor report usually says:
- •Brand visibility up 22 points
- •Coverage across ChatGPT, Perplexity and Gemini
- •40+ prompts monitored
- •Month-over-month trend line
What the same report needs to say to be readable:
- •22 points against a stated baseline and a stated band
- •Prompts per engine, runs per prompt, and the window
- •Mentions, citations and recommendations counted separately
- •Which specific source pages changed, and on which date
Get the report your current vendor cannot produce
We run your buyer prompts to a panel size that can actually resolve movement, publish the settings we used, and name the source pages behind every change. First findings inside 14 days.
Book a Discovery CallWhat our 90,132-answer corpus says the ceiling actually is
Sampling discipline tells you whether a number is readable. It does not tell you how much movement is available to buy in the first place. For that we can use our own data.
Our CITE Index study ran from May 19 to July 21, 2026: 90,132 AI answers over 63 consecutive days across 10 consumer categories. The full figures are on our AI search statistics page and in the final report. Three findings bear directly on what a service can promise.
Some categories simply do not hand out the position being sold
Across the whole study, the category leader changed on only 18.7% of day pairs. In 4 of the 10 categories, the leader never changed once in 63 days.
That is the uncomfortable number for this industry. A service selling "we will make you the answer" in a category with an entrenched incumbent is selling something the category did not hand out even once in two months of daily observation.
It does not mean the work is pointless there. It means the goal has to be second and third position, category-adjacent prompts, or a longer horizon, and that should be in the proposal rather than discovered in month four.
Ask what moved. If the answer is a percentage and not a source, nothing moved.
Winning looks like 78.2%, not 100%
The average category leader appeared in 78.2% of its own category's answers. Even dominant brands were absent from roughly one answer in five.
Any proposal that treats total absence as the problem and total presence as the fix has mis-specified the target. The realistic ask is share, and share has a known shape.
Most of the work sits on pages you do not own
Of the 12 most-cited domains in our corpus, 4 were brand-owned. Reddit was the single most-cited source in the study, appearing in 13.6% of all answers with 14,698 citations.
This is why on-site-only retainers underperform their proposals. You can rewrite every page you control and still be reading from a source pool where two thirds of the shelf belongs to someone else.
You cannot buy your way into a source pool you were never eligible for.
A managed GEO agency earns its fee mostly on that off-domain half, which is also the half that takes longest and is hardest to fake in a dashboard. If you want the full checklist for that conversation, we wrote one on how to vet a GEO agency.
The 5-step test to run before you sign
Everything above is diagnosis. Here is the procedure. It takes about two hours of a vendor's time and one email thread, and it works on any provider in the category.
Step 1: Ask for panel size before you ask for results
Request four numbers in writing: prompts tracked, engines covered, runs per prompt, and reporting window. Anything below 50 prompts per engine cannot support a monthly claim, because the 10 to 50 step is where the precision actually gets bought.
Step 2: Ask them to separate mentions from citations from recommendations
Request the same case study broken into the three events. If the vendor cannot split them retroactively, the underlying data was never stored that way, which tells you what the dashboard is doing.
Step 3: Ask which cadence they run per engine, and why
The correct answer references the difference in source turnover between engines. A vendor running one weekly cadence across ChatGPT and Perplexity and reporting both in one column has not thought about it.
Step 4: Ask for a no-change baseline period before optimization starts
Two to four weeks of measurement with no work performed establishes your category's own noise floor. Every later claim gets compared against that floor rather than against zero. A vendor who resists this is protecting the size of their month-three number.
Step 5: Ask which specific source pages they intend to change
Not "we will improve your content." Which URLs, on which domains, owned by whom. Our corpus says most of the answer is built from pages the brand does not control, so a plan that names only your own site has already capped itself.
What a real engagement looks like by month
Timelines in this category get quoted loosely. Here is the shape we see when the work is sound and the measurement is honest, alongside the number that should be reported at each stage.
| Window | What is actually happening | The number that should be reported |
|---|---|---|
| Weeks 1 to 4 | Baseline measurement with no optimization, prompt set built and frozen, source pool mapped | Your noise floor, stated as a band and not a point |
| Weeks 5 to 12 | Answer blocks deployed, entity data reconciled, off-domain placement work starts | Named pages changed, with dates. No coverage claim yet |
| Months 4 to 6 | Retrieval shifts become visible on the slower-turnover engines first | Citation share per engine, against the week 1 to 4 band |
| Months 6 to 12 | Position consolidates or the category proves locked | Recommendation rate, reported separately from mentions |
The pattern that should worry you is a large coverage claim in month two. That is the window where the work has been done but the engines have not finished re-reading the web, and a big number arriving early is usually the panel, not the brand.
Pricing for this sits in a fairly narrow band across the market, which we broke down in what AI visibility actually costs. The variable worth negotiating is panel size, not headcount.
Have your current vendor's numbers checked against a real baseline
We re-run your existing prompt set at a panel size that can resolve movement, compare it to a no-change period, and tell you which reported wins survive. No contract required to see the first pass.
Book a Discovery CallFAQ
Do AI SEO services actually work?
They work on the mechanism and usually fail on the proof. Structured answer content, consistent entity data and off-domain placement do change which sources AI engines retrieve, and that effect is well documented. The common failure is that the panel used to report it is too small to separate the effect from normal variation, so the buyer cannot tell which part of the improvement was real.
How long do AI search engine optimization services take to show results?
Plan on four to six months before a coverage claim means anything, with the first four weeks spent measuring nothing at all. The baseline period is not padding. Without it, every later number is compared against zero rather than against your category's own movement, and your category may move 13 points on its own.
How much do AI search optimization services cost?
Published retainers in this category commonly run from roughly $2,500 to $10,000 a month, with audits quoted separately. Price compares poorly across vendors because tiers meter different things. Normalize on prompts tracked per engine per month before comparing, since that is the variable that decides whether anything they report is readable.
What is the difference between AI SEO services and regular SEO services?
Regular SEO competes for positions on a results page that exists and can be checked. AI search optimization competes for inclusion in an answer that is generated once and then gone, drawn from a source pool that is mostly not your own domain. We covered the full list of deliverables in what AI SEO services include.
Why do two AI visibility tools give different numbers for the same brand?
Because each one runs a different prompt set, against a different engine mix, a different number of times, and scores it with a different counting rule. There is no reference measurement to check either against. We took that apart in detail in why LLM visibility tools disagree.
The bottom line
The vendor who sent that twelve-prompt case study was not running a scam. They were running the work most of this category runs, and reporting it with the instrument most of this category uses.
We asked them to re-run the same prompt set at 50 prompts per engine with a three-week no-change period first. They agreed. The baseline came back with a 14-point band, which meant their original 22-point claim survived, barely, as an 8-point real gain on one engine and nothing detectable on the other.
That is a worse-sounding number and a far better piece of information. It named the engine, it named the size, and it could be checked next quarter.
Ask for the settings before you ask for the results. If the settings cannot support the results, you have learned everything you needed to know without spending a month finding out. And if they can, you have found a vendor worth arguing with about the actual work, which is where the argument should have been the whole time.
Continue the brief
What Are Generative Engine Optimization Services?
Generative engine optimization services get your brand cited by ChatGPT, Perplexity, and AI Overviews. Here is what they include and how to buy.
How Do You Choose an AEO Agency in 2026?
Every AEO agency shortlist on page one was written by an AEO agency that ranked itself first. Here is what to check instead, and what the work can move.
What Does a Mature AEO Program Look Like in 2026?
Conductor surveyed 250+ enterprise marketing leaders. 51% run AEO on integrated platforms, and high-maturity teams are 6x more likely to. Here is the bar.
Framework
Learn the CITE framework behind our GEO and AEO work
See how Comprehend, Influence, Track, and Evolve turn AI visibility into an operating system.
Services
Explore our managed GEO services and AEO execution model
Audit, prompt discovery, content execution, and ongoing monitoring tied to AI search outcomes.
Audit
Start with an AI visibility audit before execution
Understand prompt coverage, recommendation gaps, source mix, and where competitors are winning.
