How do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity
Run a frozen weekly prompt panel on ChatGPT, Claude, Gemini, and Perplexity. Log mention, citation, and recommend as three columns — never one vendor score.
William Spurlock Founder — Spurlock Studios 24 MIN
You measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity by running the same frozen buyer prompts on each product, on a calendar, and logging three outcomes per run: mention, citation, and recommend. Those are different events. Semrush’s 2026 definition of AI visibility is how often a brand is mentioned, cited, or recommended in generated answers — three verbs, not one dashboard tile. There is no public console that reports “you were visible 41% this week” across those four UIs, and a vendor composite that averages them hides the engine that actually lost.
This spoke is the per-engine panel under the Answer Engine Optimization playbook. Inclusion as a concept lives in what AI visibility is. Rank tracking plus Search Console generative AI reports live in measuring AI search visibility. This page does not replace those. It tells you what to write down when a buyer is inside ChatGPT, Claude, Gemini, or Perplexity.
I have been SEO certified since 2021. The AEO version of that work is still a sheet with engine columns, not a score you cannot audit.
The short answer
- Freeze 15–40 money prompts, a competitor set, and a three-column rubric before you change a page.
- Run that panel weekly on ChatGPT, Claude, Gemini, and Perplexity. Same IDs. Same weekday. Same search / login state.
- Mention = the brand string appears. Citation = a URL, footnote, source chip, or Sources list credits your domain. Recommend = you are a pick or shortlist member on a buying prompt.
- Score each engine on its own row. Never average the four into one “AI visibility %.”
- Trend buckets over weeks. One sitting is an anecdote. Two cycles with frozen rules is a measurement.
| Column | Counts as yes | Does not count as yes |
|---|---|---|
| Mention | Brand, legal name, or agreed alias in the answer body | Your domain in a source list with no name in the prose |
| Citation | Visible credit to your URL or a clearly attributed source page on your domain | A competitor roundup that names you with their URL as the source |
| Recommend | Presented as a pick, hire, or “use them for X” on a comparison / buy prompt | Named as a discarded example, or named on a definition prompt |
A blended 40% is how a ChatGPT shortlist win funds a Gemini blog nobody asked for.
Why does one vendor score fail on four engines?
Because the products do not expose the same evidence, do not retrieve on the same cadence, and do not mean the same thing when they stay silent. A single number has to pretend those differences are noise. They are the job.
| Engine | What the UI actually shows | What a blended score pretends |
|---|---|---|
| ChatGPT | Inline citations when search ran, plus a Sources panel of consulted links (OpenAI Help Center) | Search-off memory answers equal Search-on citations |
| Claude | Inline source links when web search ran; many prompts never search | A training-only answer is a citation loss |
| Gemini | Source chips in the Gemini app; not the same surface as AI Overviews or AI Mode (Google AI features) | A Gemini mention is an Overview impression |
| Perplexity | Numbered citations to live URLs on Search answers (Perplexity Help Center) | Citation order is an authority score |
Vendor tools can still sit next to the log. They cannot replace the log unless they show the prompt list, the engine, the date, and the raw answer. If the tool cannot produce a row you could have written by hand, it is a vibe with a login.
- Keep four engine series, not one sparkline.
- Keep three outcome columns, not a weighted “visibility index.”
- Keep a
search_invokedflag on ChatGPT, Claude, and Gemini. Perplexity Search is retrieval-first; still record the mode (Search vs a Pages/shopping surface). - Keep GSC generative AI impressions on the Google Search spoke. Do not paste them into a ChatGPT cell.
If leadership only wants one number, they are asking you to hide an engine. Say that out loud.
What is mention vs citation vs recommend on these products?
What AI visibility is defines mention, citation, and share of voice. This panel adds recommend as a third operator column because buying prompts are why you are in the room. Semrush names the trio — mentioned, cited, or recommended — then tracks mentions and citations as separate metrics. Do the same. Do not invent a “healthy” citation percentage. There is no published industry target that transfers to your category.
| Outcome | Buyer sees | Ticket if you lose |
|---|---|---|
| Mention, no citation | Your name, no path | Entity packet, third-party corroboration |
| Citation, no mention | A URL with no brand shout | On-page brand string in the extractable answer |
| Recommend, no citation | A pick without a click | Quoteable comparison table on a live URL |
| Citation + recommend | A pick with a path | Defend the page; do not rewrite for sport |
| Wrong mention | False price, geo, or offer | Fact repair on every public profile that still says the old thing |
Semrush’s ghost-citation study found that in their 2026 dataset, most AI citations did not also name the brand in the answer. Treat that as evidence the columns diverge — not as a benchmark you should “beat.” Your category will not match their mix of prompts and engines.
Operator test on one answer:
- Delete every link, footnote, and source chip. What still names you is a mention.
- What disappeared was the citation surface. Log the URL if it was yours.
- On a hire / best-of / alternative prompt, were you a pick or just present? That is recommend.
- If the facts about you are wrong, log
accuracy: faileven if every other column is yes.
Share of voice can wait until the three columns exist. A ratio on collapsed events is how you fund the wrong page.
How do I freeze a weekly panel all four engines share?
Freeze the object before you freeze the calendar. A weekly ritual on a moving prompt list is not a panel. It is shopping for a better screenshot.
| Freeze this | How | Breaks the series if you… |
|---|---|---|
| Prompt IDs | P01…P40 with exact strings | Add “better” wording mid-quarter |
| Prompt job | Hire, best-of, alternative, compare, brand-probe | Mix brand-recognition prompts into the recommend rate |
| Competitor set | 3–6 names sales will defend | Swap the loser when the chart looks bad |
| Engines | ChatGPT, Claude, Gemini app, Perplexity Search | Quietly add Copilot or AI Mode into the same average |
| Locale / language | One market per series | Run US one week and UK the next |
| Scoring guide | The three columns plus accuracy | Let a new intern “use judgment” |
Do this once, version it (panel_v1), print the version on every chart:
- Interview sales for the questions that precede a call. Write them as the buyer would type them, not as keywords.
- Split branded probes (
Is {you} a fit for {job}?) into their own series. They test accuracy, not category recommend. - Exclude prompts that contain your brand from the recommend rate. “Tell me about Acme” is a recognition test.
- Cap the shared weekly set at a size one person will actually finish. 16 prompts × 4 engines is 64 rows. 40 × 4 is theater unless you staff it.
- Store the strings in the sheet, not in a Slack thread.
| Cadence | What you run | What you do not claim |
|---|---|---|
| Weekly | Money subset (hire, best-of, alternative) | A market census |
| Monthly | Full panel, including compare and brand-probes | That monthly equals weekly |
| After a ship | Only the IDs mapped to that URL | That one green cell is a program |
The freeze is the product. The run is sampling.
What columns belong on every four-engine row?
If the sheet cannot produce this table, you do not have a four-engine panel. You have four screenshots.
| Column | Type | Required | Notes |
|---|---|---|---|
date | ISO date | yes | The run day, not the day you typed the sheet |
panel | v1 / v1.1 | yes | Printed on the slide |
prompt_id | P07 | yes | Never the prose as the primary key |
engine | enum | yes | chatgpt / claude / gemini / perplexity |
mode | enum | yes | Search on/off, Gemini app, Perplexity Search |
search_invoked | yes / no / n/a | yes | n/a only if the product always retrieves |
mention | yes / no | yes | Body text only |
citation | yes / no / n/a | yes | n/a when no search ran |
cited_url | URL or empty | if citation yes | Exact path |
recommend | yes / soft / no / n/a | yes | n/a on definitional and branded-probe series |
accuracy | pass / fail / n/a | yes | n/a if you were never described |
competitors | list | yes | Spelling as the model wrote it |
artifact | path | yes | Screenshot or pasted answer + source list |
ticket | URL or none | yes | Empty ticket = theater |
Copy-once scoring rules, also frozen:
- Mention looks at the answer body, not the browser tab title, not a source card that never got spoken.
- Citation looks at visible credit to your domain. A ghost citation (URL, no name) is
citation: yes,mention: no. - Recommend is only scored on hire / best-of / alternative / compare prompts. Soft = named in a pile with no pick language.
- Accuracy fail beats a pretty recommend. Wrong price is a fail even if they shortlisted you.
- Do not compute a row-level “visibility score.” The row is the three columns.
| Prompt job | Recommend scored? | Why |
|---|---|---|
| Hire / best-of / alternative | yes | This is the money series |
| You vs rival | yes | A pick can be implied by the comparison frame |
| What is {category} | no (n/a) | Definitional; citation may still matter |
| Is {you} a fit | no (n/a) | Accuracy series, not category share |
If a vendor export cannot map onto these columns, keep the vendor in an appendix. Manage to the sheet.
How do I score ChatGPT without mixing Search and memory?
ChatGPT with search off is a memory product. ChatGPT with search on is a retrieval product. OpenAI’s Help Center is explicit: responses that use search may include inline citations; if those are missing, open Sources under the response for cited sources and other relevant links. OpenAI also says search results and citations can be incomplete, outdated, or wrong — open the source. That is your accuracy column, not a footnote you skip.
On the API side, OpenAI’s web search guide splits citations (URLs attached to the answer) from sources (the fuller list of URLs consulted). The source list is often longer than the citation list. If you sample via API, log both. If you sample in ChatGPT, log inline cites and the Sources panel as two fields, then set citation: yes only when your domain appears in either.
| Field | What you write | Trap |
|---|---|---|
search_on | yes / no (forced or observed) | Scoring search-off as a citation miss |
inline_cite | your URL or none | Counting a hover that was a competitor |
sources_panel | domains listed | Treating “other relevant links” as equal to an inline cite without labeling them |
mention | yes / no | Counting a user-uploaded file that named you |
recommend | yes / soft / no / n/a | Calling a definitional answer a pick |
Procedure for one ChatGPT row:
- New chat. Same plan you will use every week. Logged-out or a dedicated research account — pick one and freeze it.
- Set search the way the series requires. If buyers use Search, the money series is Search on.
- Paste the frozen string. No preamble, no “act as.”
- Screenshot the answer, the inline cites, and the Sources panel.
- Fill mention / citation / recommend / accuracy / competitor names / cited URLs.
Contamination table — freeze one side and stay there:
| State | What it does to the row | Freeze as |
|---|---|---|
| Logged-in with chat history | Prior threads leak into “who should I hire” | Dedicated research account, or logged-out |
| Custom GPT / project files | Your PDF named you; the open web did not | Exclude from the public-web series |
| Search off | Memory / training residue only | Separate series, or force Search on |
| Follow-up in the same thread | New query, not a repeat of P07 | Save the first answer first |
Do not use chatgpt.com referrals in analytics as the ChatGPT scoreboard. Referrals are lagging clicks after a citation already happened. Absence in GA4 is not absence in the answer.
How do I score Claude when it often does not search?
Claude will answer many prompts from prior knowledge and never open the web. Anthropic’s web search tool docs describe the retrieval path: when search runs, citations are on, and each citation carries a URL, title, and a short cited snippet. That is the API contract. The consumer UI is the same idea with different chrome — inline source links when a search happened, silence when it did not.
A 2026 capture study by Josh Blyskal observed Claude invoking search on 36.6% of their mixed prompt set (web search enabled; the rest answered without retrieval). That is one study’s mix, not a product SLA. Use it as a warning: if you zero-fill citation on every Claude row, you will “lose” weeks where Claude never searched.
| Claude state | How you know | How you score |
|---|---|---|
| Search invoked | UI shows search activity; source links appear | Mention / citation / recommend as usual |
| No search | Answer, no sources, no search indicator | search_invoked: no. Citation is n/a, not no |
| Search invoked, you absent | Sources exist, none are you | Citation no. That is a real miss |
| Brand probe, invented fact | Wrong year, offer, or geo | Accuracy fail, even with a mention |
Log these Claude-specific flags:
-
search_invoked(yes / no) -
cited_url(only when search ran) -
mention(training residue still counts as a mention) -
recommend(only on buy / compare prompts) -
accuracy(wrong is worse than absent)
Third-party writeups often claim Claude’s citations overlap heavily with Brave’s top results. Blyskal’s own paper treats Brave overlap as an observed match, not proof of a documented backend. Do not brief leadership that “ranking on Brave is the Claude ranking factor.” Log what Claude showed. If you also track Brave, keep it in a notes column.
If Claude names you with no search, that is memory / training residue. Useful. Different ticket than a Perplexity numbered miss.
How do I score Gemini without treating it as AI Overviews?
Gemini the assistant is not Google AI Overviews and it is not AI Mode. Google’s AI features documentation covers Overviews and AI Mode as Search surfaces. Those two “may use different models and techniques, so the set of responses and links they show will vary.” The Gemini app is a third place a buyer can sit. Put it on this panel. Put Overview / AI Mode impressions on measuring AI search visibility via Search Console generative AI reports.
If you sample Gemini via API with Search grounding, Google’s grounding docs return groundingMetadata with source URLs when grounding actually ran. Grounding can fail to attach. Consumer Gemini shows source chips when it used the web. Log the product you actually ran. Do not join API grounding chunks to an app screenshot and call it one rate.
| Surface | Where the buyer is | Where you log it |
|---|---|---|
| Gemini app | gemini.google.com / Gemini apps | This four-engine panel |
| AI Overviews | Classic Search | GSC gen-AI + Overview prompt log |
| AI Mode | Conversational Search | Separate AI Mode log; Google says links can differ from Overviews |
| Gemini API + Search grounding | Your app, not the consumer assistant | Own series, labeled API |
Gemini row checklist:
- Same frozen string, new conversation, frozen login state.
- Record whether source chips appeared (
grounded: yes/no). - Mention / citation / recommend / accuracy as usual.
- Write the cited domains. A YouTube or Reddit chip that names you is a mention path, not an owned citation, unless your domain is the chip.
- Never copy a GSC impression count into the Gemini mention cell.
Signed-in Google accounts personalize. Freeze a research login or a signed-out window the same way you freeze ChatGPT. Do not mix a Workspace account that has your site in Drive with a clean consumer run and call the delta “visibility.”
An Overview screenshot is not a Gemini win. A Gemini source chip is not an AI Mode citation.
How do I score Perplexity’s numbered sources?
Perplexity Search is the cleanest citation UI of the four. Perplexity’s Help Center describes real-time web search with numbered citations to original sources so a person can verify. That numbered list is the citation object. The prose is the mention / recommend object. Do not mix Search with Pages or shopping. Those are different products and a different scoreboard.
| Perplexity field | Yes means | Common lie |
|---|---|---|
| Mention | Brand in the answer body | Counting a source title that is not spoken in the prose |
| Citation | Your domain in the numbered list | Treating citation number as a published rank |
| Recommend | You are a pick in the prose on a buy prompt | A numbered cite to your docs on a definitional “what is” query |
| Source URL | The exact path | Collapsing every path to the homepage and calling it one page that “works” |
Perplexity procedure:
- Use Search, not a Collections/Pages experiment, unless that is a labeled series.
- Paste the frozen prompt. No follow-up until the row is saved — follow-ups are new queries.
- Copy the numbered source list in order. Order is presentation, not an official authority score.
- Mark mention, citation, recommend, accuracy, competitors.
- If you were retrieved in a product that exposes a fuller result set but not cited, log retrieved-vs-cited only when you actually have that field. The consumer UI usually shows cites, not the discarded retrieval set.
Do not publish a Perplexity “citation rate target.” Perplexity does not. Your week-over-week rate on frozen IDs is the number. A blog that tells you 80% is “healthy” is selling a round number.
What does one money prompt look like across four engines?
Work one hire prompt through the four UIs on the same day. Copy the columns. Do not invent a composite.
Prompt P07 (example shape, not a client result): “Who should I hire for {category} for a {ICP} team that already has {constraint}?”
| Engine | Search / grounded | Mention | Citation | Recommend | What you write in notes |
|---|---|---|---|---|---|
| ChatGPT | Search on | Yes | /services in Sources | Yes | Inline cite was a third-party roundup; Sources also listed you |
| Claude | No search | Yes | n/a | Soft | Named in a long list, no sources, no pick language |
| Gemini | Source chips | No | Competitor docs | No | Three chips, none yours |
| Perplexity | Numbered cites | Yes | Homepage [3] | No | Named as an also-ran; pick was Competitor A with [1] |
That single prompt is four stories:
| Reading | Wrong response | Right ticket |
|---|---|---|
| ChatGPT recommend + weak owned cite | “We won ChatGPT” | Make the owned comparison URL the cite, not a roundup |
| Claude mention, no search | “Claude hates us” | Keep the residue; do not treat n/a citation as a miss |
| Gemini absence | “Google is broken” | Fix Gemini-app extractability; check Overviews separately |
| Perplexity mention, rival cited first | “We need more blogs” | A citeable table the numbered list can lift |
If you average those four rows into 50%, you will brief a number that never happened on any engine.
Save artifacts per row:
- Timestamp and panel version
- Engine and mode
- Full answer text or screenshot
- Source list as copied, not as remembered
- Competitor names exactly as spelled
- Ticket owner, or
noneif nothing ships
A row without an artifact is a rumor.
How do I keep the weekly freeze honest?
The weekly part is the part teams skip. They run a heroic Tuesday, screenshot the green cells, and come back in six weeks with different wording. That is not a freeze. That is two anecdotes.
| Rule | Freeze | Allowed change |
|---|---|---|
| Prompt text | Exact string | New IDs in panel_v2, never silent edits to P07 |
| Weekday | Same local day, Eastern if you are this studio | Shift the whole series; do not mix Tue and Fri in one chart |
| Account | Logged-out / dedicated research login | Personal Plus history is a different product |
| Model picker | Whatever the consumer default is, recorded | A one-off “try the other model” row, labeled, excluded from the rate |
| Follow-ups | None until the row is saved | A labeled follow-up series if buyers actually chat that way |
| Location | One geo | A second series for a second market |
Honesty checklist before you hit send on the deck:
- Panel version on the slide matches the sheet
- Engine filters are visible. No stacked bar that sums ChatGPT with Perplexity
- Search-off and search-on are not in the same citation rate
- Recommend rate excludes branded probes
- Accuracy fails are listed, not averaged away
- n/a Claude citations are not counted in the denominator as losses
- No “healthy range” band copied from a vendor blog
Same prompt, different day, different sources. That variance is why you trend. It is also why one heroic run is worthless.
If nobody will run it next Tuesday, do not start. Assign an owner or skip the theater.
What usually fails first on a four-engine panel?
The first failure is collapsing the three columns. The second is collapsing the four engines. The third is treating a no-search Claude answer as a Perplexity-style citation miss. Those three errors show up before anyone has a crawl problem — though crawl still kills the retrieval engines when someone pasted a “block AI” robots snippet.
| Failure | What it looks like | Cost | Do this instead |
|---|---|---|---|
| One vendor score | “AI visibility 22%” | Funds the engine that is already fine | Four rates, three columns |
| Citation % target | “We need to hit 40% cites” | Rewrites that chase a made-up band | Trend your own frozen IDs |
| Search-off mixed in | ChatGPT memory week vs Search week | Fake drops | search_on as a required field |
| Gemini = Overviews | GSC line pasted into Gemini | Misses the app the buyer used | Separate surfaces |
| Prompt shopping | New wording after a bad week | Un-auditable lift | Version the panel |
| No artifacts | Sheet of yes/no with no screenshot | Arguments in the QBR | Answer + sources saved |
| GA4 as the panel | “ChatGPT traffic is down” | Optimizing for clicks you never measured in the answer | Panel first, referrers lag |
Crawl still matters on the retrieval path. If ChatGPT Search, Perplexity, and Gemini all go dark the week a WAF started challenging AI user-agents, start with fetch, not with new outlines. That eligibility work is documented on the playbook’s crawl spokes. This panel will show the outage. It will not explain robots.txt by itself.
The expensive failure is a sparkline with no owner. A KPI that cannot become a ticket is decoration.
How do I turn a four-engine miss into a ticket?
A miss that does not name a URL or a corroboration target will die in the deck. Map the hole to one object.
| Engine × column hole | Likely object | Not the ticket |
|---|---|---|
| ChatGPT mention no, Search on | Third-party roundups and entity packet | A new homepage slogan |
| ChatGPT recommend yes, owned cite no | Comparison / pricing URL the Sources panel can lift | More brand-name blog posts |
| Claude no-search mention wrong | About + every profile that still has the old year | A Brave ranking campaign you cannot audit |
| Claude search on, citation no | Quoteable primary page + fetch for Claude’s search/user agents | Treating n/a weeks as this miss |
| Gemini app mention no | Extractable answer in HTML, not a client-only shell | Pasting GSC Overview impressions here |
| Perplexity mention yes, citation no | Unique numbers and a table on a live URL | Buying a Pages experiment |
| Perplexity citation yes, recommend no | Pick language and a shortlist table a numbered cite can lift | Changing the prompt until you are first |
| All three retrieval engines dark the same week | robots.txt / WAF / 403s | A content calendar |
Ticket hygiene:
- One miss → one owner → one URL or one corroboration ask.
- Re-run only the mapped prompt IDs the week the URL ships.
- If the cell does not move after the ship, the post did not change inclusion. It changed the CMS.
Do not open a ticket for a Claude no-search citation: n/a. That row is not a miss.
What should I skip if I only have a week?
Skip SOV math, skip a 40-prompt fantasy, skip buying a tracker before you have columns, skip blending Google Search into this sheet.
| Do in week one | Skip |
|---|---|
| 12–16 hire / best-of / alternative prompts | Branded vanity prompts in the recommend rate |
| Four engines, one weekday, artifacts | Copilot, AI Mode, Overviews in the same average |
| Mention / citation / recommend / accuracy | Sentiment scores you will not act on |
| One competitor set sales accepts | Industry “benchmark” citation % |
| Map each miss to a URL or a corroboration target | A content calendar of 20 outlines |
Order for a short week:
- Write the scoring guide in the sheet header. Three columns,
search_invoked,n/arules for Claude. - Lock 12 strings. Print them. Do not “improve” them on Friday.
- Run all four engines on the same day. 48 rows. That is a baseline, not a program.
- Pick three tickets from the worst engine × column pair. Ship those, not a blog about measurement.
- Put next Tuesday on a calendar with a named owner.
If you cannot finish 12 × 4, cut prompts, not engines. Four engines on eight prompts beats two engines on thirty you will abandon.
A week cannot prove a trend. It can prove you have a log. If week two does not exist on the calendar before week one starts, you are collecting souvenirs.
How do I report four engines without averaging them?
Bring a table, not a pie. The pie is how a Perplexity citation week hides a ChatGPT recommend collapse.
Slide object
- Panel
v1, dates, market, frozen competitors. - Counts first: n of N money prompts, per engine, for mention, citation, and recommend.
- The hole: engine × column that is actually losing.
- Three tickets with URLs.
- Ask: keep the weekly freeze, or book a visibility audit if nobody owns the log.
| Exec question | Bad answer | Better answer |
|---|---|---|
| “What’s our AI visibility?” | “34%” | “ChatGPT Search: 6/16 recommends, 3/16 owned cites. Claude: 4/16 mentions, 9/16 no-search. Gemini app: 2/16 mentions. Perplexity: 8/16 cites, 3/16 recommends. Panel v1, week of [date].” |
| “Are we healthy?” | A vendor band | “No portable healthy %. Here is three weeks on the same IDs.” |
| “Which engine first?” | The one with a LinkedIn screenshot | Deal-influence × current hole |
| “Did the rewrite work?” | Sessions | Same IDs, same engines, two cycles later |
| “Can we use Semrush’s number?” | Yes, silently | Yes, and the prompt list, or we manage to the sheet |
If a vendor number goes in the appendix, quote that product’s formula and engine coverage. Semrush Brand Performance is not this DIY event-count. Mixing them across quarters is how a “decline” appears that is really a formula change.
Do not put GSC generative AI impressions on this slide as if they were ChatGPT. Point to the measurement spoke for Google Search. Keep this slide for the four chat/answer UIs.
Four small multiples beat one average. Always.
When is a four-engine panel not worth running yet?
Skip it when the inputs are fiction, when buyers never use these products, or when nobody will sample twice. A weekly freeze with no second week is a workshop.
| Skip this panel when | Do this first | Then start |
|---|---|---|
| You cannot name 12 buyer prompts | A one-hour sales interview | Lock panel_v1 |
| Money URLs do not fetch | Crawl, WAF, noindex accidents | After five URLs return in view-source |
| Legal blocks every fetcher | A written allow-list for citation bots | Retrieval engines only, labeled |
| You are mid-rebrand | Freeze the legal name on About | After entity pages agree |
| Buyers only use Google Search | Rank + GSC gen-AI + Overview log | Add these four if a deal cites ChatGPT |
| Leadership wants one number and will not take four | Refuse the average | Audit, or no program |
| No owner for next Tuesday | Assign or stop | Week 1 and week 3 on the calendar |
DIY this if one named person will finish the weekly subset. Hire when the four engines disagree, the argument is political, and you need an outside prioritization hammer. Lane work is /visibility. The system map is the AEO playbook.
A visibility audit is the right next step when you cannot produce mention vs citation vs recommend, split by engine, after two honest weeks — not when you dislike the first percentages.
Do not wait for a vendor to invent a score you could have written down on Tuesday.
FAQ
How do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity?
Run the same frozen buyer prompts on each product every week and log mention, citation, and recommend as separate columns, plus whether search or grounding actually ran. Score each engine on its own row. There is no official four-engine console and no transferable “healthy” citation rate — your trend on frozen IDs is the measurement.
How do I measure whether do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity is working?
It is working when next week’s sheet uses the same prompt IDs, the same scoring guide, and artifacts you can reopen — and when a moved cell becomes a ticket on a URL. If the only output is a blended percentage, or if Claude no-search rows are scored as citation losses, the ritual is not working. Two cycles minimum before you brief a lift.
What usually fails first when teams try this?
They collapse mention, citation, and recommend into one score, then average the four engines. Next they mix ChatGPT Search-off with Search-on, and they treat a Claude answer that never searched as a Perplexity-style miss. Fix the scoreboard before you buy another tracker. Crawl failures still matter, but they show up as retrieval engines going dark together, not as a mysterious composite dip.
How long does this take to show results?
A baseline is one working day for 12–16 prompts across four engines. A trend takes at least two weekly cycles on frozen IDs — longer if you only sample monthly. Fetch and entity-packet fixes can move retrieval engines within weeks; they are not a citation SLA. Do not declare a win from one Tuesday.
What should I skip if I only have a week?
Skip share-of-voice math, branded vanity prompts inside the recommend rate, and any blend with AI Overviews or Copilot. Lock 12 strings, run all four engines in one day, save artifacts, and pick three tickets from the worst engine × column pair. Cut prompts before you cut engines if time runs out.
When is this not worth doing yet?
When you cannot name money prompts, when public URLs do not fetch, when buyers never use these four products, or when nobody will run the panel twice. Also skip it if leadership will only accept one blended number — that request is the opposite of this scoreboard. Keep Google measurement on its own spoke until a deal actually cites ChatGPT, Claude, Gemini, or Perplexity.
CTA
If you cannot see mention, citation, and recommend on ChatGPT, Claude, Gemini, and Perplexity as four logs, you cannot manage them. Freeze the panel. Stop averaging engines.
Lane: /visibility · Next step: visibility audit
What questions does this article answer?
- How do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity?
- Run the same frozen buyer prompts on each product every week and log mention, citation, and recommend as separate columns, plus whether search or grounding actually ran. Score each engine on its own row. There is no official four-engine console and no transferable “healthy” citation rate — your trend on frozen IDs is the measurement.
- How do I measure whether do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity is working?
- It is working when next week’s sheet uses the same prompt IDs, the same scoring guide, and artifacts you can reopen — and when a moved cell becomes a ticket on a URL. If the only output is a blended percentage, or if Claude no-search rows are scored as citation losses, the ritual is not working. Two cycles minimum before you brief a lift.
- What usually fails first when teams try this?
- They collapse mention, citation, and recommend into one score, then average the four engines. Next they mix ChatGPT Search-off with Search-on, and they treat a Claude answer that never searched as a Perplexity-style miss. Fix the scoreboard before you buy another tracker. Crawl failures still matter, but they show up as retrieval engines going dark together, not as a mysterious composite dip.
- How long does this take to show results?
- A baseline is one working day for 12–16 prompts across four engines. A trend takes at least two weekly cycles on frozen IDs — longer if you only sample monthly. Fetch and entity-packet fixes can move retrieval engines within weeks; they are not a citation SLA. Do not declare a win from one Tuesday.
- What should I skip if I only have a week?
- Skip share-of-voice math, branded vanity prompts inside the recommend rate, and any blend with AI Overviews or Copilot. Lock 12 strings, run all four engines in one day, save artifacts, and pick three tickets from the worst engine × column pair. Cut prompts before you cut engines if time runs out.
- When is this not worth doing yet?
- When you cannot name money prompts, when public URLs do not fetch, when buyers never use these four products, or when nobody will run the panel twice. Also skip it if leadership will only accept one blended number — that request is the opposite of this scoreboard. Keep Google measurement on its own spoke until a deal actually cites ChatGPT, Claude, Gemini, or Perplexity.
Last reviewed — OpenAI ChatGPT Search source UI, Anthropic web-search citations, Google AI-features docs, Perplexity Help Center citations, Semrush mention/cite/recommend definition, and Josh Blyskal 2026 Claude search-invocation study checked 2026-09-05.
AI Visibility
AI Visibility Cannabis visibility when the ad accounts are banned
Google and Meta will not take the usual spend. The models still answer dispensary, cultivator, and brand questions — if the site can be read and the cart can clear a 21+ order.
AI Visibility When ChatGPT names the franchise, not your shop
Run the best-HVAC-near-me prompt panel. If the model names a national franchise, fix corroboration and entity facts — not another blog calendar.
AI Visibility How do I get cited by Perplexity specifically
Allow PerplexityBot, put a liftable answer and unique numbers in HTML, then log numbered sources on a frozen prompt panel. There is no bought citation rate.
AI Visibility What belongs in an AI visibility monthly retainer vs a one-time audit
A one-time audit is the baseline plus prioritized fixes. A monthly retainer is prompt-panel tracking, entity hygiene, page jobs, and citation recovery.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.