AEO 101Single source of truth on AEO
AI Visibility12 min read

What Does Profound AI Actually Measure?

Subia Peerzada

Subia Peerzada

Founder, Cite Solutions · July 29, 2026

Every review of Profound AI you will find compares the same things: engine count, feature list, price. All three are on the vendor's own pricing page, which makes most of those reviews a slower way to read it.

The question nobody answers is the one that decides whether the subscription works. When the dashboard says your share of voice fell from 14% to 9%, did your visibility fall, or did you just draw a different sample from the same distribution?

That is not a rhetorical question. It has an arithmetic answer, and you can compute it from Profound's published limits before you ever open a demo.

Sampling resolution by tier

Both self-serve tiers buy the same 30 answers per prompt, per engine

Published Profound plan limits divided by prompts and engines, set against the 33 to 94 answer convergence range measured across 30 platform-topic combinations in arxiv 2607.10341.

30

answers per prompt per engine per month, on both Starter and Growth

The lowest convergence threshold measured in the paper was 33. Three of its 30 test cases never stabilized at all, even after 125 answers.

Starter · $99/mo

30 / prompt / engine / mo

1 engine50 prompts1,500 responses/mo

One engine, ChatGPT only. 50 prompts against 1,500 monthly responses.

Growth · $399/mo

30 / prompt / engine / mo

3 engines100 prompts9,000 responses/mo

Three engines. Six times the responses, spread across three times the surfaces.

Enterprise · Custom

quoted / prompt / engine / mo

9 enginesprompts quotedresponses tailored

Up to nine engines. Response allowance is quoted, so the resolution is negotiable.

Why waiting three months does not fix it

Thirty answers a month becomes 90 by the end of a quarter, which clears the floor. That only counts if the thing being measured held still for all three months. ChatGPT shipped GPT-5.5 in May and GPT-5.6 on July 9, and Semrush measured only 25.6% overlap in cited domains between two reasoning depths on the same prompt.

The samples have to land inside a window where the model has not changed. In 2026 that window is weeks.

Sources: Profound published pricing (tryprofound.com/pricing, July 2026) · arxiv 2607.10341, Sielinski, July 11 2026 · Semrush reasoning-mode study, June 30 2026. Per-prompt figures are our own arithmetic on the published limits.

What does Profound AI actually measure?

Profound AI measures how often your brand appears and gets cited in answers from up to nine AI engines, including ChatGPT, Perplexity, Gemini, Claude, Copilot, and Google AI Overviews. It runs a fixed prompt set on a schedule, then scores mentions, citation sources, sentiment, and competitor share. It measures the answers it sampled, not the answers your buyers received.

That last sentence is the whole post. Everything below is the consequence.

Profound is not selling you a number. It is selling you a sample.

What Profound measures well, and what no dashboard can measure

Profound is the strongest instrument in this category, and the reason is data nobody else has. It is worth being precise about where that strength actually sits.

It measures a real prompt distribution, not a list somebody guessed

Most AI visibility tools track prompts a marketer typed into a setup screen. Profound's index is built on more than 1.5 billion real-user prompts, which means its prompt-volume data reflects questions people actually asked rather than questions a growth team imagined.

That distinction matters more than engine count. A perfectly sampled answer to a prompt nobody types is still zero information.

It measures the source pool beside you, which is the part you can act on

The citation-source panel is the most operationally useful screen in the product. It tells you which domains the model leaned on to build the answer, and that list is your off-page target map.

Our concluded CITE Index study of 90,132 AI answers found Reddit was the single most-cited source, appearing 14,698 times across 13.6% of all answers. Four of the twelve most-cited domains were brand-owned sites. A source panel tells you which half of that split you are competing in.

It measures depth on commerce surfaces almost nobody else covers

Profound's ChatGPT Shopping teardown analyzed 812,190 product cards across 201,137 prompt runs collected in one week of June 2026. That is a real research operation, not a content-marketing exercise, and the shopping surface is genuinely under-instrumented elsewhere.

It cannot measure whether the number it gave you is stable

This is not a Profound flaw. No prompt-tracking platform reports the confidence interval around its own score, because doing so would make most weekly movement look like what it is.

What a feature comparison tells you:

  • How many engines are covered
  • Whether citation share is broken out from mentions
  • Whether the export format fits your reporting stack
  • What it costs per seat

What a sampling audit tells you:

  • How many answers produced each score
  • Whether that count clears the threshold where rankings stabilize
  • Whether the samples landed inside one model generation
  • Whether last week's move was signal or resampling

Every published review of Profound answers the first list. None of them answer the second.

What each Profound pricing tier buys you in sample size

Here is the arithmetic. Profound publishes prompt limits, response allowances, and engine counts on its pricing page. Divide the second by the first two and you get the only number that governs reliability.

TierPriceEnginesPromptsResponses/moAnswers per prompt, per engine, per month
Starter$99/mo billed yearly1 (ChatGPT)501,50030
Growth$399/mo billed yearly31009,00030
EnterpriseQuotedUp to 9TailoredTailoredWhatever you negotiate

Growth costs four times Starter and buys six times the responses. It also spreads them across three times the surfaces and twice the prompts. The resolution per prompt per engine is identical.

The threshold that number needs to clear

On July 11, 2026, Ronald Sielinski of IQRush published From Stochastic to Stable, a convergence framework for deciding when an AI visibility measurement has collected enough answers to be trusted. Across 30 platform-topic combinations on Gemini, SearchGPT, and Perplexity, stable rankings required between 33 and 94 answers.

Three of the 30 test cases never stabilized, even after 125 answers. Search Engine Journal's coverage framed the finding bluntly, and a separate reproduction by researchers at the University of St. Gallen reached the same place.

Thirty answers sits below the low end of that range.

Two honest caveats before you quote this at a vendor

The paper is an unreviewed preprint, its questions were generated by ChatGPT rather than drawn from real searches, and Sielinski says plainly that the exact numbers will not transfer cleanly to other topics. The method transfers. The 33 does not.

So treat 33 as an order of magnitude, not a certification line. The point is not that 30 fails and 34 passes. The point is that 30 is the same order as the threshold, which means the reliability of your dashboard is a live question rather than a settled one.

Waiting a quarter does not solve it

Thirty answers a month becomes 90 across a quarter, which clears the floor comfortably. That only counts if the process being measured held still for three months.

It did not. ChatGPT moved to GPT-5.5 in May and shipped GPT-5.6 on July 9. Semrush found that the same prompt at two reasoning depths returns cited-domain sets that overlap only 25.6%, with Reddit's share halving from 15% to 7% and government and academic sources rising from 1.9% to 8.8%.

Pooling three months of answers across two model generations does not give you a bigger sample. It gives you a blended average of two different systems.

A dashboard reading is not a measurement until you know how many samples produced it.

Before you buy a tier, find out how many samples your category actually needs

We run your buyer prompts to convergence across every engine, report the band rather than the point estimate, and tell you which movements in your current dashboard were real. First findings inside 14 days.

Book a Discovery Call

6 reasons your Profound dashboard moved when your visibility did not

None of these are bugs. Every one is a property of measuring a generative system with a finite budget, and every one shows up as a red or green arrow that looks like news.

Reason #1: You resampled, and the resample landed differently

At roughly one answer per prompt per engine per day, a single week gives you seven readings. A share-of-voice figure built on seven answers will wobble by several points on its own.

This is the most common false alarm in the category, and it is indistinguishable from a real change unless you are tracking the spread as well as the mean.

Reason #2: The engine changed its reasoning depth on you

Semrush's finding is the sharpest version of this. Three quarters of the cited domains turn over between minimal and high reasoning on an identical prompt. Any dashboard that does not record which mode produced each answer is averaging two different systems and labeling the result "ChatGPT."

Reason #3: A model shipped, and nobody told your dashboard

Model releases do not come with a changelog entry for your brand. GPT-5.6 landed on July 9, 2026. The share your reports showed on July 8 and July 10 were measurements of two different products.

Reason #4: Your prompt set mixes intents that behave differently

Conductor ran 14,000 API calls across 10 industries, seven intent types, four models, and five personas. Brand overlap between repeated runs ranged from 40% on purchase intent to 63% on comparison intent.

Intent, not industry, predicted consistency. A blended score across mixed-intent prompts is arithmetic on incompatible units.

Reason #5: You are watching the middle of the ranking, where the noise lives

Our own data argues against the strongest version of the noise thesis, and it is worth saying so. Across 63 days and 90,132 answers, the category leader appeared in 78.2% of its category's answers, the leader flipped on only 18.7% of day pairs, and in four of ten categories the leader never changed at all.

Top-of-category positions are far more stable than the volatility narrative suggests. The churn is concentrated in positions four through ten, which is exactly where most B2B brands sit and exactly where a 30-answer sample has the least to say.

Reason #6: You are reading a rank when you should be reading a rate

Ranks amplify small differences. Two brands separated by half a percentage point of citation share can swap positions on a single answer, and the dashboard will render that as a position change rather than as a coin flip.

Sielinski's illustration makes the point: a gap of roughly 9.5% against 6.0% citation share disappears entirely below the convergence threshold.

The score did not move because your visibility moved. It moved because you sampled again.

Step 1: Write down the decision the dashboard is supposed to inform

Before you compare tiers, write one sentence naming the decision you will make differently based on what the tool reports. "Whether to invest in review acquisition this quarter" is a decision. "Understanding our AI visibility" is not.

A tier that cannot resolve the difference you would act on is the wrong tier, no matter what it costs.

Step 2: Size the prompt set to convergence, not to the tier limit

Take your ten highest-intent buyer prompts and run each one 40 to 50 times on a single engine before you commit to anything. Plot how the ranking moves as answers accumulate and stop when it flattens.

That number is your category's real convergence point. Most teams discover they need fewer prompts sampled far more deeply, which is the opposite of how every tier is packaged.

Step 3: Segment every score by engine, intent, and reasoning mode

Never report a blended cross-engine number. Our study found ChatGPT cited a source in 92.5% of its answers, Google AI Mode in 97.4%, and Gemini in 79.1%. Averaging engines with an 18-point spread in citation behavior produces a figure that describes none of them.

Split intent the same way, on Conductor's evidence, and record reasoning mode on every run.

Step 4: Report a rate with a band, never a rank

Replace "we are number three" with "we appeared in 31% of answers, plus or minus 9 points, on 30 samples." It is a less satisfying slide and a far more honest one.

The band is computable from your sample size. Once it is on the chart, the weekly arrows that used to generate meetings stop generating them, which is the point.

Step 5: Pair the instrument with the work that changes the reading

A measurement platform tells you where you stand. It does not write the answer block, fix the passage the model could not extract, or earn the third-party mention that puts you in the source pool.

That gap is the whole reason the managed GEO agency model exists alongside the tools. If you have an owner who will act on the data every week, buy the platform. If you do not, the subscription becomes a tab nobody opens.

Buy the instrument. Do not mistake it for the work.

When Profound is the right call, and when it is not

We recommend Profound regularly. Being clear about the fit is more useful than being clear about the price.

SituationVerdictWhy
Enterprise brand, dedicated owner, nine engines matterStrong fitDeepest prompt-volume dataset in the category, plus SSO, SOC 2, and API access at the quoted tier.
Consumer or commerce brand tracking ChatGPT ShoppingStrong fitThe shopping surface is barely instrumented anywhere else.
B2B team wanting one defensible number per quarterNegotiate the response allowanceEngine coverage is not your constraint. Answers per prompt is.
Series A startup, one marketer, no analyst timePoor fitStarter's single engine and 30 answers per prompt will not resolve the differences you would act on.
Nobody has been named as the weekly ownerWrong purchase entirelyThe measurement is not the bottleneck. The follow-through is.

If you are still shortlisting, our AI visibility platform buyer's guide covers the six jobs any serious platform should do, and how to choose AI visibility tools covers the lighter end of the market. For head-to-head reads, see Profound vs Otterly and Profound vs AthenaHQ, or the full field in our guide to GEO tools for 2026.

Where the two measurements disagree, our breakdown of GEO versus SEO covers why the same page can rank well on Google and never appear in an AI answer.

FAQ

What does Profound AI do?

Profound AI runs a fixed set of prompts against AI engines on a schedule and reports how your brand appears in the answers. Its named modules cover Answer Engine Insights, Prompt Volumes, Agent Analytics, Shopping, and Agents. It tracks up to nine engines at the enterprise tier, including ChatGPT, Perplexity, Gemini, Claude, Grok, Copilot, Meta AI, DeepSeek, and Google AI Overviews.

How much does Profound AI cost?

Profound publishes two self-serve tiers and one quoted tier. Starter is $99 per month billed yearly for one engine, 50 prompts, and 1,500 monthly responses. Growth is $399 per month billed yearly for three engines, 100 prompts, and 9,000 monthly responses. Enterprise is custom-quoted and adds up to nine engines, SSO, SOC 2, API access, and ChatGPT Shopping.

Is Profound AI worth it?

It is worth it when three things are true: you have named an owner who acts on the data weekly, your buyers use more than one engine, and you negotiate enough monthly responses to sample each prompt deeply rather than broadly. It is poor value when bought as a quarterly reporting artifact, because at self-serve response limits a single prompt gets about 30 answers per engine per month, which is around the threshold where published research says rankings begin to stabilize.

What are the best Profound AI alternatives?

The named alternatives in the category include Peec AI, Scrunch AI, Otterly, Semrush AI Visibility, and Ahrefs Brand Radar. Compare them on answers per prompt per engine rather than on engine count, because sampling depth is what determines whether the score can be trusted. We publish head-to-head reads on Profound against Otterly and against AthenaHQ, plus a full survey of the GEO tooling field for 2026.

Is Profound SEO or GEO?

Neither label fits cleanly. Profound measures generative engine optimization outcomes, meaning citations and mentions inside AI answers, rather than positions on a results page. Traditional SEO tools measure rankings, and the two rarely agree. The same page can win one and lose the other, which is why the two measurements belong on separate lines of a report.

The bottom line

Profound is a good instrument. The reason to be careful is not the product, it is the habit the product encourages, which is treating a point estimate as a fact because it arrived on a chart.

Ask any vendor how many answers produced each number, insist on rates with bands instead of ranks, and negotiate response allowance before engine count. If the answer to "how many samples is this" makes the salesperson uncomfortable, you have learned something more useful than the demo was going to teach you.

Then go do the work the dashboard points at. Nothing in the software does that part, and an AI visibility audit will tell you where the gap actually is before you commit to a year of anything.

Get a measurement you can defend in a board meeting

Cite Solutions samples your buyer prompts to convergence across every major AI engine, reports confidence bands instead of vanity ranks, and runs the content and off-page work that moves the number.

Book a Discovery Call

Ready to become the answer AI gives?

Book a 30-minute discovery call. We'll show you what AI says about your brand today. No pitch. Just data.

.md