MCP

New

AI Visibility: A Practical Framework for Getting Cited in AI Search

Written By

Written By

Chetan Parmar

Chetan Parmar

Published on

Published on

GEO ≠ SEO: A hard technical look at the evidence

AI Visibility: A Practical Framework for Getting Cited in AI Search

Your founder forwards a screenshot of ChatGPT recommending three competitors in your category, none of them you, and asks what the plan is. The reflex answer is more content. That reflex is why most AI visibility work stalls by month three: you publish twelve posts, the screenshot looks the same in April, and now the budget line has no defence.

The problem is not that the team is publishing the wrong things. It is that the program has no denominator. Nobody named an owner, nobody wrote down the baseline before the work started, and nobody agreed on what a win looks like in a system where the output changes between two runs of the same prompt. So the reporting defaults to what is countable, which is posts shipped. And posts shipped is exactly the number that does not answer the founder's question.

What follows is the program design we would set up at the start: an owner, a weekly loop, three lanes of work, and a budget line denominated in citations rather than output. It is a method, not a tactics list. Built so that in week four you can put a page in front of a founder and have the conversation end.

Budget the program against citations earned, not posts published

Budget and report an AI visibility program against citations earned per tracked prompt set. The retrieval systems you are trying to influence select passages and sources. They have no view of how many posts you published. Everything after this depends on that.

That distinction changes the arithmetic. If you publish twelve assets in a quarter and two of them ever surface as a cited source in an AI answer, your real cost per cited asset is six times your blended content cost. Most teams never calculate that number, because their reporting layer counts the twelve and stops.

The unit to budget against is a citation event: one instance where an AI answer to a prompt in your tracked set references a specific URL, yours or somebody else's. Count those by month, by platform, and by which lane of work produced them. Your budget conversation then becomes a cost per citation event, and cost per citation is a number a founder can argue with, approve, or cut on rational grounds.

This is also the honest framing for a category where nobody can promise a number. You cannot commit to a citation count next quarter. You can commit to a measured baseline, a rate of change, and a decision rule for what gets funded again.

Name the owner and the weekly loop before you shortlist a tool

Tool selection comes third. Before that, decide who owns the number, and decide what happens every week.

Retrieval is not stable. The same prompt run twice can return different sources, models re-crawl on their own schedule, and platform behaviour shifts without notice. A team that checks weekly builds a series, and a series is what lets you tell a real change from variance.

The loop has five moves, and one person is accountable for all of them. Pull the current answers for your tracked prompt set. Diff against last week. Decide which changes are worth acting on. Ship the work. Re-measure the specific thing you changed, not the whole program.

The owner should be a named person on your team, ideally whoever already owns organic. Splitting AI search visibility across content, SEO, and PR without an owner produces three partial views and no decision. When the founder asks why the number moved, somebody has to answer in a sentence, and committees do not.

Set the AI visibility baseline your founder will ask you about in week one

Before you change anything, write down the current state. This takes about two days, and it is the artefact that does the most work in the program, because every later claim of improvement is measured against it.

Build a prompt set of roughly 40 to 60 queries. Split it deliberately: branded prompts where your name is in the query, category prompts where a buyer is asking for options, problem prompts where the buyer describes a pain and no vendor is named, and comparison prompts where two competitors are named and you are not. Run each prompt several times per platform, across ChatGPT, Google AI Overview, Gemini and Claude. One run of a system that varies between runs tells you almost nothing.

For each run, record four things: whether your brand appeared at all, which domains were cited, whether any of those domains were yours, and where in the answer your mention sat. Mention rate across repeated runs is the baseline. A screenshot cannot stand in for it.

The failure mode here is a flattering baseline. If your prompt set is mostly branded queries, your mention rate will look strong and tell you nothing, because a model asked directly about your company will usually find your own site. The non-branded prompts are where the money is, and where the number will be uncomfortable. Keep both, report both separately, and never blend them into one score.

Split the work into three lanes: pages you own, sources you can earn, answers you must correct

Once the baseline exists, sort every candidate action into one of three lanes. Lanes matter because they have different owners, different lead times, and different failure modes. Mix them into one backlog and the slowest lane silently starves the fastest.

Lane one is pages you own. Your site, your docs, your comparison pages, your pricing explanation. You control the text and can ship in days. The work is structural: answer the question in the passage rather than building to it, put the specific claim near the heading that predicts it, and make each page resolve one question completely rather than three questions partially.

Lane two is sources you can earn. Third-party listicles, review platforms, community threads, industry roundups, partner documentation, podcast transcripts. You do not control these and cannot ship them on your own schedule. Mechanically, a model answering "best X for Y" has no particular reason to prefer a vendor's own page over a third-party list, which puts a large share of category-query citations on domains you do not own.

That is a mechanism argument, not a measurement, and your own prompt set will settle it for your category. It is also the lane in-house teams tend to under-resource, for the same reason: nothing in it ships when you want it to.

Lane three is answers you must correct. Wrong pricing, a stale positioning line, an acquisition the model has not registered, a category label that puts you next to the wrong competitors. Correcting these means fixing the source the model is drawing from, not writing a new blog post and hoping.

Lane

What actually moves it

Typical lead time

How you verify

Pages you own

Passage structure, page scope, internal linking

Days to weeks

Your URL appears as a cited source for the target prompt

Sources you can earn

Outreach, review programs, community participation, analyst and partner content

Weeks to a quarter

A third-party URL that mentions you appears as a cited source

Answers you must correct

Fixing the underlying source, publishing an unambiguous canonical statement

Unpredictable

The incorrect claim stops appearing across repeated runs

Watch the middle row. It has the longest lead time and usually the highest ceiling, which is exactly why it gets cut first when a quarter goes sideways.

Decide what a monitoring tool can and cannot structurally do inside this loop

Tools in this category, including Profound, Peec AI, Scrunch, Otterly and Ahrefs Brand Radar, are built to sample prompts at scale and report what the answers said. That is real work. Doing it manually across four platforms and sixty prompts every week is not sustainable for a five-person marketing team.

What a monitoring layer can do: run the prompt set on a schedule, hold the history so you see a trend rather than a snapshot, break results out by platform, and tell you which domains are getting cited in your category.

What it cannot do, no matter how good the interface is: decide which lane a given gap belongs to, write the third-party placement, run the outreach that gets you onto the roundup, or make the judgement call that a competitor's sudden lift came from a review-site refresh rather than their blog. Those are decisions and execution. A reporting surface does not make decisions. That gap is what we built LLMLab around, which is why the analysis in our product resolves into a task with an owner rather than a chart.

Be honest with yourself about which half your team is short on. If you have five marketers and no visibility data, buy visibility. If you have data and nothing has moved in two quarters, the loop is what you are missing, and another dashboard will not supply it.

Report AI visibility in a format a founder reads in 90 seconds

The monthly report is one page. If it needs a walkthrough, it will not survive a board week.

Put four things on it: share of answers on the non-branded prompt set, by platform, with last month next to it; citation events this month, split by the three lanes; the domains most cited in your category answers, with a yes or no on whether you appear on each; and a line describing the single largest change you made, plus what happened to the specific prompts that change targeted.

That last line does most of the work. It turns the report from a status update into a record of cause and effect, and it is what lets you argue for the budget again. A founder who can see that a review-platform refresh moved seven prompts will fund the next one. A founder who sees "published nine posts, LLM visibility up slightly" will ask what those posts had to do with it, and you will not have an answer.

This sounds like it only works if you already have domain authority

Partly true, and the honest version deserves precision. Retrieval leans on sources that are already crawled, linked and trusted, so a two-year-old domain does not walk into a head query and displace an incumbent's documentation this quarter. If your category has an entrenched reference source that every answer leans on, you are not replacing it with better formatting.

Where the objection breaks down is the earnable lane. The citations you are competing for on comparison and problem queries frequently sit on third-party domains that already have the authority, and getting listed there is an outreach and review problem rather than a link-equity problem. A challenger brand cannot outrank a market leader's docs. It does not need to, if it appears in the roundups the model retrieves instead.

The second thing authority does not settle is specificity. A page that answers one narrow question completely, with a number, a source and a date, gives a retrieval system something quotable. A high-authority page that hedges for four hundred words before saying anything gives it nothing to lift. That is reasoning from how retrieval works rather than a measured claim, and you should treat it as such. It is also testable inside your own prompt set within a month.

Retire any lane that has not moved citation volume in a quarter

Programs die from accumulation. Three lanes turn into seven workstreams, each defended by whoever started it, and by month nine nobody can say which one is producing.

Give every lane one full quarter. The earn lane cannot show results in six weeks, and killing it early is the most common self-inflicted wound in this work. At the quarter mark, compare citation events on your tracked prompt set against hours spent, and ignore traffic, impressions and posts shipped.

Then make one of three calls per lane: fund it again at the same level, fund it harder because the cost per citation is the lowest of the three, or stop it and move the hours. Write the decision down with the number that justified it. That written record is what stops a retired lane from quietly restarting in two quarters because someone new joined and had an idea.

One failure mode to watch: retiring a lane when the real problem was execution quality inside it. If the earn lane produced nothing and you sent nine generic outreach emails, the problem sits in the work rather than the lane.

Run the first four weeks of the loop

Four weeks is enough to have the whole program standing and one real result to point at.

  1. Week one, baseline. Build the 40 to 60 prompt set, run it across all four platforms with repeated runs, and record mention rate, cited domains and your rank position. Name the owner in the same week.

  2. Week two, sort and ship the fast lane. Assign every gap to one of the three lanes. Ship two or three page-level fixes on the prompts where a competitor is cited and you have an equivalent page that simply answers less directly.

  3. Week three, start the slow lane. Identify the five third-party domains cited most often in your category answers, and start the review, outreach or contribution process for each. Fix any factually wrong statement about your company at its source.

  4. Week four, re-measure and report. Re-run the prompt set, diff against week one, and build the one-page report. Include the changes that produced nothing, because that is the credibility of the whole document.

Do not expect the earn lane to have landed by week four. Expect the loop to be running, the baseline to be defensible, and the founder conversation to be about cost per citation instead of about a screenshot.

If you want to see what your own baseline looks like before committing a quarter to this, we run a free discovery call and walk through the numbers with you. No pitch, and we do these for people who are still deciding whether the category is real.

We are LLMLab, also written LLM Labs. We help B2B brands get recommended by AI assistants across ChatGPT, Google AI Overview, Gemini and Claude.

FAQs

What counts as a citation event in AI search?

A citation event is one instance where an AI answer to a prompt in your tracked set references a specific URL as a source. Count it per prompt, per platform, and per run, since the same prompt can return different sources on repeat runs. Separate citation events where the URL is yours from ones where a third-party page that mentions you was cited, because those come from different lanes of work.

How often should you re-measure AI search visibility?

Weekly for the tracked prompt set, monthly for the report that leaves your team. The weekly cadence exists to separate real movement from run-to-run variance, not because you should act on every fluctuation. Anything less frequent than monthly and you lose the ability to attribute a change to something you did.

Does schema markup improve LLM visibility?

Structured data helps machines parse what a page asserts, and it is cheap to implement, so there is no reason to skip it. There is no public confirmation from the major assistant providers that schema is a direct ranking or selection input for generative answers, so treat it as hygiene rather than a lever. Judge it the way you judge every other change: by whether the prompts it targeted moved.

How many prompts should a tracked set contain?

Roughly 40 to 60 is a workable starting point for a single product line, weighted toward non-branded category, problem and comparison queries. Fewer than about 30 and normal variance swamps your signal. If you sell into several distinct segments, build a separate set per segment rather than stretching one set thin.

How long does it take for changes to show up in AI answers?

Page-level changes on your own domain can surface within days to a few weeks, depending on how quickly the page is recrawled and indexed. Third-party placements move on the publisher's schedule, which is often a quarter or longer. Nobody can promise a timeline here, which is precisely why the program needs a baseline and a weekly series rather than a target date.

AI Visibility: A Practical Framework for Getting Cited in AI Search

Your founder forwards a screenshot of ChatGPT recommending three competitors in your category, none of them you, and asks what the plan is. The reflex answer is more content. That reflex is why most AI visibility work stalls by month three: you publish twelve posts, the screenshot looks the same in April, and now the budget line has no defence.

The problem is not that the team is publishing the wrong things. It is that the program has no denominator. Nobody named an owner, nobody wrote down the baseline before the work started, and nobody agreed on what a win looks like in a system where the output changes between two runs of the same prompt. So the reporting defaults to what is countable, which is posts shipped. And posts shipped is exactly the number that does not answer the founder's question.

What follows is the program design we would set up at the start: an owner, a weekly loop, three lanes of work, and a budget line denominated in citations rather than output. It is a method, not a tactics list. Built so that in week four you can put a page in front of a founder and have the conversation end.

Budget the program against citations earned, not posts published

Budget and report an AI visibility program against citations earned per tracked prompt set. The retrieval systems you are trying to influence select passages and sources. They have no view of how many posts you published. Everything after this depends on that.

That distinction changes the arithmetic. If you publish twelve assets in a quarter and two of them ever surface as a cited source in an AI answer, your real cost per cited asset is six times your blended content cost. Most teams never calculate that number, because their reporting layer counts the twelve and stops.

The unit to budget against is a citation event: one instance where an AI answer to a prompt in your tracked set references a specific URL, yours or somebody else's. Count those by month, by platform, and by which lane of work produced them. Your budget conversation then becomes a cost per citation event, and cost per citation is a number a founder can argue with, approve, or cut on rational grounds.

This is also the honest framing for a category where nobody can promise a number. You cannot commit to a citation count next quarter. You can commit to a measured baseline, a rate of change, and a decision rule for what gets funded again.

Name the owner and the weekly loop before you shortlist a tool

Tool selection comes third. Before that, decide who owns the number, and decide what happens every week.

Retrieval is not stable. The same prompt run twice can return different sources, models re-crawl on their own schedule, and platform behaviour shifts without notice. A team that checks weekly builds a series, and a series is what lets you tell a real change from variance.

The loop has five moves, and one person is accountable for all of them. Pull the current answers for your tracked prompt set. Diff against last week. Decide which changes are worth acting on. Ship the work. Re-measure the specific thing you changed, not the whole program.

The owner should be a named person on your team, ideally whoever already owns organic. Splitting AI search visibility across content, SEO, and PR without an owner produces three partial views and no decision. When the founder asks why the number moved, somebody has to answer in a sentence, and committees do not.

Set the AI visibility baseline your founder will ask you about in week one

Before you change anything, write down the current state. This takes about two days, and it is the artefact that does the most work in the program, because every later claim of improvement is measured against it.

Build a prompt set of roughly 40 to 60 queries. Split it deliberately: branded prompts where your name is in the query, category prompts where a buyer is asking for options, problem prompts where the buyer describes a pain and no vendor is named, and comparison prompts where two competitors are named and you are not. Run each prompt several times per platform, across ChatGPT, Google AI Overview, Gemini and Claude. One run of a system that varies between runs tells you almost nothing.

For each run, record four things: whether your brand appeared at all, which domains were cited, whether any of those domains were yours, and where in the answer your mention sat. Mention rate across repeated runs is the baseline. A screenshot cannot stand in for it.

The failure mode here is a flattering baseline. If your prompt set is mostly branded queries, your mention rate will look strong and tell you nothing, because a model asked directly about your company will usually find your own site. The non-branded prompts are where the money is, and where the number will be uncomfortable. Keep both, report both separately, and never blend them into one score.

Split the work into three lanes: pages you own, sources you can earn, answers you must correct

Once the baseline exists, sort every candidate action into one of three lanes. Lanes matter because they have different owners, different lead times, and different failure modes. Mix them into one backlog and the slowest lane silently starves the fastest.

Lane one is pages you own. Your site, your docs, your comparison pages, your pricing explanation. You control the text and can ship in days. The work is structural: answer the question in the passage rather than building to it, put the specific claim near the heading that predicts it, and make each page resolve one question completely rather than three questions partially.

Lane two is sources you can earn. Third-party listicles, review platforms, community threads, industry roundups, partner documentation, podcast transcripts. You do not control these and cannot ship them on your own schedule. Mechanically, a model answering "best X for Y" has no particular reason to prefer a vendor's own page over a third-party list, which puts a large share of category-query citations on domains you do not own.

That is a mechanism argument, not a measurement, and your own prompt set will settle it for your category. It is also the lane in-house teams tend to under-resource, for the same reason: nothing in it ships when you want it to.

Lane three is answers you must correct. Wrong pricing, a stale positioning line, an acquisition the model has not registered, a category label that puts you next to the wrong competitors. Correcting these means fixing the source the model is drawing from, not writing a new blog post and hoping.

Lane

What actually moves it

Typical lead time

How you verify

Pages you own

Passage structure, page scope, internal linking

Days to weeks

Your URL appears as a cited source for the target prompt

Sources you can earn

Outreach, review programs, community participation, analyst and partner content

Weeks to a quarter

A third-party URL that mentions you appears as a cited source

Answers you must correct

Fixing the underlying source, publishing an unambiguous canonical statement

Unpredictable

The incorrect claim stops appearing across repeated runs

Watch the middle row. It has the longest lead time and usually the highest ceiling, which is exactly why it gets cut first when a quarter goes sideways.

Decide what a monitoring tool can and cannot structurally do inside this loop

Tools in this category, including Profound, Peec AI, Scrunch, Otterly and Ahrefs Brand Radar, are built to sample prompts at scale and report what the answers said. That is real work. Doing it manually across four platforms and sixty prompts every week is not sustainable for a five-person marketing team.

What a monitoring layer can do: run the prompt set on a schedule, hold the history so you see a trend rather than a snapshot, break results out by platform, and tell you which domains are getting cited in your category.

What it cannot do, no matter how good the interface is: decide which lane a given gap belongs to, write the third-party placement, run the outreach that gets you onto the roundup, or make the judgement call that a competitor's sudden lift came from a review-site refresh rather than their blog. Those are decisions and execution. A reporting surface does not make decisions. That gap is what we built LLMLab around, which is why the analysis in our product resolves into a task with an owner rather than a chart.

Be honest with yourself about which half your team is short on. If you have five marketers and no visibility data, buy visibility. If you have data and nothing has moved in two quarters, the loop is what you are missing, and another dashboard will not supply it.

Report AI visibility in a format a founder reads in 90 seconds

The monthly report is one page. If it needs a walkthrough, it will not survive a board week.

Put four things on it: share of answers on the non-branded prompt set, by platform, with last month next to it; citation events this month, split by the three lanes; the domains most cited in your category answers, with a yes or no on whether you appear on each; and a line describing the single largest change you made, plus what happened to the specific prompts that change targeted.

That last line does most of the work. It turns the report from a status update into a record of cause and effect, and it is what lets you argue for the budget again. A founder who can see that a review-platform refresh moved seven prompts will fund the next one. A founder who sees "published nine posts, LLM visibility up slightly" will ask what those posts had to do with it, and you will not have an answer.

This sounds like it only works if you already have domain authority

Partly true, and the honest version deserves precision. Retrieval leans on sources that are already crawled, linked and trusted, so a two-year-old domain does not walk into a head query and displace an incumbent's documentation this quarter. If your category has an entrenched reference source that every answer leans on, you are not replacing it with better formatting.

Where the objection breaks down is the earnable lane. The citations you are competing for on comparison and problem queries frequently sit on third-party domains that already have the authority, and getting listed there is an outreach and review problem rather than a link-equity problem. A challenger brand cannot outrank a market leader's docs. It does not need to, if it appears in the roundups the model retrieves instead.

The second thing authority does not settle is specificity. A page that answers one narrow question completely, with a number, a source and a date, gives a retrieval system something quotable. A high-authority page that hedges for four hundred words before saying anything gives it nothing to lift. That is reasoning from how retrieval works rather than a measured claim, and you should treat it as such. It is also testable inside your own prompt set within a month.

Retire any lane that has not moved citation volume in a quarter

Programs die from accumulation. Three lanes turn into seven workstreams, each defended by whoever started it, and by month nine nobody can say which one is producing.

Give every lane one full quarter. The earn lane cannot show results in six weeks, and killing it early is the most common self-inflicted wound in this work. At the quarter mark, compare citation events on your tracked prompt set against hours spent, and ignore traffic, impressions and posts shipped.

Then make one of three calls per lane: fund it again at the same level, fund it harder because the cost per citation is the lowest of the three, or stop it and move the hours. Write the decision down with the number that justified it. That written record is what stops a retired lane from quietly restarting in two quarters because someone new joined and had an idea.

One failure mode to watch: retiring a lane when the real problem was execution quality inside it. If the earn lane produced nothing and you sent nine generic outreach emails, the problem sits in the work rather than the lane.

Run the first four weeks of the loop

Four weeks is enough to have the whole program standing and one real result to point at.

  1. Week one, baseline. Build the 40 to 60 prompt set, run it across all four platforms with repeated runs, and record mention rate, cited domains and your rank position. Name the owner in the same week.

  2. Week two, sort and ship the fast lane. Assign every gap to one of the three lanes. Ship two or three page-level fixes on the prompts where a competitor is cited and you have an equivalent page that simply answers less directly.

  3. Week three, start the slow lane. Identify the five third-party domains cited most often in your category answers, and start the review, outreach or contribution process for each. Fix any factually wrong statement about your company at its source.

  4. Week four, re-measure and report. Re-run the prompt set, diff against week one, and build the one-page report. Include the changes that produced nothing, because that is the credibility of the whole document.

Do not expect the earn lane to have landed by week four. Expect the loop to be running, the baseline to be defensible, and the founder conversation to be about cost per citation instead of about a screenshot.

If you want to see what your own baseline looks like before committing a quarter to this, we run a free discovery call and walk through the numbers with you. No pitch, and we do these for people who are still deciding whether the category is real.

We are LLMLab, also written LLM Labs. We help B2B brands get recommended by AI assistants across ChatGPT, Google AI Overview, Gemini and Claude.

FAQs

What counts as a citation event in AI search?

A citation event is one instance where an AI answer to a prompt in your tracked set references a specific URL as a source. Count it per prompt, per platform, and per run, since the same prompt can return different sources on repeat runs. Separate citation events where the URL is yours from ones where a third-party page that mentions you was cited, because those come from different lanes of work.

How often should you re-measure AI search visibility?

Weekly for the tracked prompt set, monthly for the report that leaves your team. The weekly cadence exists to separate real movement from run-to-run variance, not because you should act on every fluctuation. Anything less frequent than monthly and you lose the ability to attribute a change to something you did.

Does schema markup improve LLM visibility?

Structured data helps machines parse what a page asserts, and it is cheap to implement, so there is no reason to skip it. There is no public confirmation from the major assistant providers that schema is a direct ranking or selection input for generative answers, so treat it as hygiene rather than a lever. Judge it the way you judge every other change: by whether the prompts it targeted moved.

How many prompts should a tracked set contain?

Roughly 40 to 60 is a workable starting point for a single product line, weighted toward non-branded category, problem and comparison queries. Fewer than about 30 and normal variance swamps your signal. If you sell into several distinct segments, build a separate set per segment rather than stretching one set thin.

How long does it take for changes to show up in AI answers?

Page-level changes on your own domain can surface within days to a few weeks, depending on how quickly the page is recrawled and indexed. Third-party placements move on the publisher's schedule, which is often a quarter or longer. Nobody can promise a timeline here, which is precisely why the program needs a baseline and a weekly series rather than a target date.

Get a free AI Visibility report

Get a free AI Visibility report

Get a free AI Visibility report

Free AI Visibility report on how your brand appear in

ChatGPT and Google AI Overview

Free AI Visibility report on how your brand appear in ChatGPT and Google AI Overview

GEO ≠ SEO: A hard technical look at the evidence
GEO ≠ SEO: A hard technical look at the evidence
GEO ≠ SEO: A hard technical look at the evidence
GEO ≠ SEO: A hard technical look at the evidence