Spurlock Studios
Contact
Share LinkedIn X
A lime beam hitting a small brass nameplate. Thesis: MEASURE BRAND VISIBILITY CHATGPT CLAUDE.

You measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity by running the same frozen buyer prompts on each product, on a calendar, and logging three outcomes per run: mention, citation, and recommend. Those are different events. Semrush’s 2026 definition of AI visibility is how often a brand is mentioned, cited, or recommended in generated answers — three verbs, not one dashboard tile. There is no public console that reports “you were visible 41% this week” across those four UIs, and a vendor composite that averages them hides the engine that actually lost.

This spoke is the per-engine panel under the Answer Engine Optimization playbook. Inclusion as a concept lives in what AI visibility is. Rank tracking plus Search Console generative AI reports live in measuring AI search visibility. This page does not replace those. It tells you what to write down when a buyer is inside ChatGPT, Claude, Gemini, or Perplexity.

I have been SEO certified since 2021. The AEO version of that work is still a sheet with engine columns, not a score you cannot audit.

The short answer

  • Freeze 15–40 money prompts, a competitor set, and a three-column rubric before you change a page.
  • Run that panel weekly on ChatGPT, Claude, Gemini, and Perplexity. Same IDs. Same weekday. Same search / login state.
  • Mention = the brand string appears. Citation = a URL, footnote, source chip, or Sources list credits your domain. Recommend = you are a pick or shortlist member on a buying prompt.
  • Score each engine on its own row. Never average the four into one “AI visibility %.”
  • Trend buckets over weeks. One sitting is an anecdote. Two cycles with frozen rules is a measurement.
ColumnCounts as yesDoes not count as yes
MentionBrand, legal name, or agreed alias in the answer bodyYour domain in a source list with no name in the prose
CitationVisible credit to your URL or a clearly attributed source page on your domainA competitor roundup that names you with their URL as the source
RecommendPresented as a pick, hire, or “use them for X” on a comparison / buy promptNamed as a discarded example, or named on a definition prompt

A blended 40% is how a ChatGPT shortlist win funds a Gemini blog nobody asked for.

Why does one vendor score fail on four engines?

Because the products do not expose the same evidence, do not retrieve on the same cadence, and do not mean the same thing when they stay silent. A single number has to pretend those differences are noise. They are the job.

EngineWhat the UI actually showsWhat a blended score pretends
ChatGPTInline citations when search ran, plus a Sources panel of consulted links (OpenAI Help Center)Search-off memory answers equal Search-on citations
ClaudeInline source links when web search ran; many prompts never searchA training-only answer is a citation loss
GeminiSource chips in the Gemini app; not the same surface as AI Overviews or AI Mode (Google AI features)A Gemini mention is an Overview impression
PerplexityNumbered citations to live URLs on Search answers (Perplexity Help Center)Citation order is an authority score

Vendor tools can still sit next to the log. They cannot replace the log unless they show the prompt list, the engine, the date, and the raw answer. If the tool cannot produce a row you could have written by hand, it is a vibe with a login.

  1. Keep four engine series, not one sparkline.
  2. Keep three outcome columns, not a weighted “visibility index.”
  3. Keep a search_invoked flag on ChatGPT, Claude, and Gemini. Perplexity Search is retrieval-first; still record the mode (Search vs a Pages/shopping surface).
  4. Keep GSC generative AI impressions on the Google Search spoke. Do not paste them into a ChatGPT cell.

If leadership only wants one number, they are asking you to hide an engine. Say that out loud.

What is mention vs citation vs recommend on these products?

What AI visibility is defines mention, citation, and share of voice. This panel adds recommend as a third operator column because buying prompts are why you are in the room. Semrush names the trio — mentioned, cited, or recommended — then tracks mentions and citations as separate metrics. Do the same. Do not invent a “healthy” citation percentage. There is no published industry target that transfers to your category.

OutcomeBuyer seesTicket if you lose
Mention, no citationYour name, no pathEntity packet, third-party corroboration
Citation, no mentionA URL with no brand shoutOn-page brand string in the extractable answer
Recommend, no citationA pick without a clickQuoteable comparison table on a live URL
Citation + recommendA pick with a pathDefend the page; do not rewrite for sport
Wrong mentionFalse price, geo, or offerFact repair on every public profile that still says the old thing

Semrush’s ghost-citation study found that in their 2026 dataset, most AI citations did not also name the brand in the answer. Treat that as evidence the columns diverge — not as a benchmark you should “beat.” Your category will not match their mix of prompts and engines.

Operator test on one answer:

  • Delete every link, footnote, and source chip. What still names you is a mention.
  • What disappeared was the citation surface. Log the URL if it was yours.
  • On a hire / best-of / alternative prompt, were you a pick or just present? That is recommend.
  • If the facts about you are wrong, log accuracy: fail even if every other column is yes.

Share of voice can wait until the three columns exist. A ratio on collapsed events is how you fund the wrong page.

How do I freeze a weekly panel all four engines share?

Freeze the object before you freeze the calendar. A weekly ritual on a moving prompt list is not a panel. It is shopping for a better screenshot.

Freeze thisHowBreaks the series if you…
Prompt IDsP01…P40 with exact stringsAdd “better” wording mid-quarter
Prompt jobHire, best-of, alternative, compare, brand-probeMix brand-recognition prompts into the recommend rate
Competitor set3–6 names sales will defendSwap the loser when the chart looks bad
EnginesChatGPT, Claude, Gemini app, Perplexity SearchQuietly add Copilot or AI Mode into the same average
Locale / languageOne market per seriesRun US one week and UK the next
Scoring guideThe three columns plus accuracyLet a new intern “use judgment”

Do this once, version it (panel_v1), print the version on every chart:

  1. Interview sales for the questions that precede a call. Write them as the buyer would type them, not as keywords.
  2. Split branded probes (Is {you} a fit for {job}?) into their own series. They test accuracy, not category recommend.
  3. Exclude prompts that contain your brand from the recommend rate. “Tell me about Acme” is a recognition test.
  4. Cap the shared weekly set at a size one person will actually finish. 16 prompts × 4 engines is 64 rows. 40 × 4 is theater unless you staff it.
  5. Store the strings in the sheet, not in a Slack thread.
CadenceWhat you runWhat you do not claim
WeeklyMoney subset (hire, best-of, alternative)A market census
MonthlyFull panel, including compare and brand-probesThat monthly equals weekly
After a shipOnly the IDs mapped to that URLThat one green cell is a program

The freeze is the product. The run is sampling.

What columns belong on every four-engine row?

If the sheet cannot produce this table, you do not have a four-engine panel. You have four screenshots.

ColumnTypeRequiredNotes
dateISO dateyesThe run day, not the day you typed the sheet
panelv1 / v1.1yesPrinted on the slide
prompt_idP07yesNever the prose as the primary key
engineenumyeschatgpt / claude / gemini / perplexity
modeenumyesSearch on/off, Gemini app, Perplexity Search
search_invokedyes / no / n/ayesn/a only if the product always retrieves
mentionyes / noyesBody text only
citationyes / no / n/ayesn/a when no search ran
cited_urlURL or emptyif citation yesExact path
recommendyes / soft / no / n/ayesn/a on definitional and branded-probe series
accuracypass / fail / n/ayesn/a if you were never described
competitorslistyesSpelling as the model wrote it
artifactpathyesScreenshot or pasted answer + source list
ticketURL or noneyesEmpty ticket = theater

Copy-once scoring rules, also frozen:

  1. Mention looks at the answer body, not the browser tab title, not a source card that never got spoken.
  2. Citation looks at visible credit to your domain. A ghost citation (URL, no name) is citation: yes, mention: no.
  3. Recommend is only scored on hire / best-of / alternative / compare prompts. Soft = named in a pile with no pick language.
  4. Accuracy fail beats a pretty recommend. Wrong price is a fail even if they shortlisted you.
  5. Do not compute a row-level “visibility score.” The row is the three columns.
Prompt jobRecommend scored?Why
Hire / best-of / alternativeyesThis is the money series
You vs rivalyesA pick can be implied by the comparison frame
What is {category}no (n/a)Definitional; citation may still matter
Is {you} a fitno (n/a)Accuracy series, not category share

If a vendor export cannot map onto these columns, keep the vendor in an appendix. Manage to the sheet.

How do I score ChatGPT without mixing Search and memory?

ChatGPT with search off is a memory product. ChatGPT with search on is a retrieval product. OpenAI’s Help Center is explicit: responses that use search may include inline citations; if those are missing, open Sources under the response for cited sources and other relevant links. OpenAI also says search results and citations can be incomplete, outdated, or wrong — open the source. That is your accuracy column, not a footnote you skip.

On the API side, OpenAI’s web search guide splits citations (URLs attached to the answer) from sources (the fuller list of URLs consulted). The source list is often longer than the citation list. If you sample via API, log both. If you sample in ChatGPT, log inline cites and the Sources panel as two fields, then set citation: yes only when your domain appears in either.

FieldWhat you writeTrap
search_onyes / no (forced or observed)Scoring search-off as a citation miss
inline_citeyour URL or noneCounting a hover that was a competitor
sources_paneldomains listedTreating “other relevant links” as equal to an inline cite without labeling them
mentionyes / noCounting a user-uploaded file that named you
recommendyes / soft / no / n/aCalling a definitional answer a pick

Procedure for one ChatGPT row:

  1. New chat. Same plan you will use every week. Logged-out or a dedicated research account — pick one and freeze it.
  2. Set search the way the series requires. If buyers use Search, the money series is Search on.
  3. Paste the frozen string. No preamble, no “act as.”
  4. Screenshot the answer, the inline cites, and the Sources panel.
  5. Fill mention / citation / recommend / accuracy / competitor names / cited URLs.

Contamination table — freeze one side and stay there:

StateWhat it does to the rowFreeze as
Logged-in with chat historyPrior threads leak into “who should I hire”Dedicated research account, or logged-out
Custom GPT / project filesYour PDF named you; the open web did notExclude from the public-web series
Search offMemory / training residue onlySeparate series, or force Search on
Follow-up in the same threadNew query, not a repeat of P07Save the first answer first

Do not use chatgpt.com referrals in analytics as the ChatGPT scoreboard. Referrals are lagging clicks after a citation already happened. Absence in GA4 is not absence in the answer.

Claude will answer many prompts from prior knowledge and never open the web. Anthropic’s web search tool docs describe the retrieval path: when search runs, citations are on, and each citation carries a URL, title, and a short cited snippet. That is the API contract. The consumer UI is the same idea with different chrome — inline source links when a search happened, silence when it did not.

A 2026 capture study by Josh Blyskal observed Claude invoking search on 36.6% of their mixed prompt set (web search enabled; the rest answered without retrieval). That is one study’s mix, not a product SLA. Use it as a warning: if you zero-fill citation on every Claude row, you will “lose” weeks where Claude never searched.

Claude stateHow you knowHow you score
Search invokedUI shows search activity; source links appearMention / citation / recommend as usual
No searchAnswer, no sources, no search indicatorsearch_invoked: no. Citation is n/a, not no
Search invoked, you absentSources exist, none are youCitation no. That is a real miss
Brand probe, invented factWrong year, offer, or geoAccuracy fail, even with a mention

Log these Claude-specific flags:

  • search_invoked (yes / no)
  • cited_url (only when search ran)
  • mention (training residue still counts as a mention)
  • recommend (only on buy / compare prompts)
  • accuracy (wrong is worse than absent)

Third-party writeups often claim Claude’s citations overlap heavily with Brave’s top results. Blyskal’s own paper treats Brave overlap as an observed match, not proof of a documented backend. Do not brief leadership that “ranking on Brave is the Claude ranking factor.” Log what Claude showed. If you also track Brave, keep it in a notes column.

If Claude names you with no search, that is memory / training residue. Useful. Different ticket than a Perplexity numbered miss.

How do I score Gemini without treating it as AI Overviews?

Gemini the assistant is not Google AI Overviews and it is not AI Mode. Google’s AI features documentation covers Overviews and AI Mode as Search surfaces. Those two “may use different models and techniques, so the set of responses and links they show will vary.” The Gemini app is a third place a buyer can sit. Put it on this panel. Put Overview / AI Mode impressions on measuring AI search visibility via Search Console generative AI reports.

If you sample Gemini via API with Search grounding, Google’s grounding docs return groundingMetadata with source URLs when grounding actually ran. Grounding can fail to attach. Consumer Gemini shows source chips when it used the web. Log the product you actually ran. Do not join API grounding chunks to an app screenshot and call it one rate.

SurfaceWhere the buyer isWhere you log it
Gemini appgemini.google.com / Gemini appsThis four-engine panel
AI OverviewsClassic SearchGSC gen-AI + Overview prompt log
AI ModeConversational SearchSeparate AI Mode log; Google says links can differ from Overviews
Gemini API + Search groundingYour app, not the consumer assistantOwn series, labeled API

Gemini row checklist:

  1. Same frozen string, new conversation, frozen login state.
  2. Record whether source chips appeared (grounded: yes/no).
  3. Mention / citation / recommend / accuracy as usual.
  4. Write the cited domains. A YouTube or Reddit chip that names you is a mention path, not an owned citation, unless your domain is the chip.
  5. Never copy a GSC impression count into the Gemini mention cell.

Signed-in Google accounts personalize. Freeze a research login or a signed-out window the same way you freeze ChatGPT. Do not mix a Workspace account that has your site in Drive with a clean consumer run and call the delta “visibility.”

An Overview screenshot is not a Gemini win. A Gemini source chip is not an AI Mode citation.

How do I score Perplexity’s numbered sources?

Perplexity Search is the cleanest citation UI of the four. Perplexity’s Help Center describes real-time web search with numbered citations to original sources so a person can verify. That numbered list is the citation object. The prose is the mention / recommend object. Do not mix Search with Pages or shopping. Those are different products and a different scoreboard.

Perplexity fieldYes meansCommon lie
MentionBrand in the answer bodyCounting a source title that is not spoken in the prose
CitationYour domain in the numbered listTreating citation number as a published rank
RecommendYou are a pick in the prose on a buy promptA numbered cite to your docs on a definitional “what is” query
Source URLThe exact pathCollapsing every path to the homepage and calling it one page that “works”

Perplexity procedure:

  1. Use Search, not a Collections/Pages experiment, unless that is a labeled series.
  2. Paste the frozen prompt. No follow-up until the row is saved — follow-ups are new queries.
  3. Copy the numbered source list in order. Order is presentation, not an official authority score.
  4. Mark mention, citation, recommend, accuracy, competitors.
  5. If you were retrieved in a product that exposes a fuller result set but not cited, log retrieved-vs-cited only when you actually have that field. The consumer UI usually shows cites, not the discarded retrieval set.

Do not publish a Perplexity “citation rate target.” Perplexity does not. Your week-over-week rate on frozen IDs is the number. A blog that tells you 80% is “healthy” is selling a round number.

What does one money prompt look like across four engines?

Work one hire prompt through the four UIs on the same day. Copy the columns. Do not invent a composite.

Prompt P07 (example shape, not a client result): “Who should I hire for {category} for a {ICP} team that already has {constraint}?”

EngineSearch / groundedMentionCitationRecommendWhat you write in notes
ChatGPTSearch onYes/services in SourcesYesInline cite was a third-party roundup; Sources also listed you
ClaudeNo searchYesn/aSoftNamed in a long list, no sources, no pick language
GeminiSource chipsNoCompetitor docsNoThree chips, none yours
PerplexityNumbered citesYesHomepage [3]NoNamed as an also-ran; pick was Competitor A with [1]

That single prompt is four stories:

ReadingWrong responseRight ticket
ChatGPT recommend + weak owned cite“We won ChatGPT”Make the owned comparison URL the cite, not a roundup
Claude mention, no search“Claude hates us”Keep the residue; do not treat n/a citation as a miss
Gemini absence“Google is broken”Fix Gemini-app extractability; check Overviews separately
Perplexity mention, rival cited first“We need more blogs”A citeable table the numbered list can lift

If you average those four rows into 50%, you will brief a number that never happened on any engine.

Save artifacts per row:

  • Timestamp and panel version
  • Engine and mode
  • Full answer text or screenshot
  • Source list as copied, not as remembered
  • Competitor names exactly as spelled
  • Ticket owner, or none if nothing ships

A row without an artifact is a rumor.

How do I keep the weekly freeze honest?

The weekly part is the part teams skip. They run a heroic Tuesday, screenshot the green cells, and come back in six weeks with different wording. That is not a freeze. That is two anecdotes.

RuleFreezeAllowed change
Prompt textExact stringNew IDs in panel_v2, never silent edits to P07
WeekdaySame local day, Eastern if you are this studioShift the whole series; do not mix Tue and Fri in one chart
AccountLogged-out / dedicated research loginPersonal Plus history is a different product
Model pickerWhatever the consumer default is, recordedA one-off “try the other model” row, labeled, excluded from the rate
Follow-upsNone until the row is savedA labeled follow-up series if buyers actually chat that way
LocationOne geoA second series for a second market

Honesty checklist before you hit send on the deck:

  • Panel version on the slide matches the sheet
  • Engine filters are visible. No stacked bar that sums ChatGPT with Perplexity
  • Search-off and search-on are not in the same citation rate
  • Recommend rate excludes branded probes
  • Accuracy fails are listed, not averaged away
  • n/a Claude citations are not counted in the denominator as losses
  • No “healthy range” band copied from a vendor blog

Same prompt, different day, different sources. That variance is why you trend. It is also why one heroic run is worthless.

If nobody will run it next Tuesday, do not start. Assign an owner or skip the theater.

What usually fails first on a four-engine panel?

The first failure is collapsing the three columns. The second is collapsing the four engines. The third is treating a no-search Claude answer as a Perplexity-style citation miss. Those three errors show up before anyone has a crawl problem — though crawl still kills the retrieval engines when someone pasted a “block AI” robots snippet.

FailureWhat it looks likeCostDo this instead
One vendor score“AI visibility 22%”Funds the engine that is already fineFour rates, three columns
Citation % target“We need to hit 40% cites”Rewrites that chase a made-up bandTrend your own frozen IDs
Search-off mixed inChatGPT memory week vs Search weekFake dropssearch_on as a required field
Gemini = OverviewsGSC line pasted into GeminiMisses the app the buyer usedSeparate surfaces
Prompt shoppingNew wording after a bad weekUn-auditable liftVersion the panel
No artifactsSheet of yes/no with no screenshotArguments in the QBRAnswer + sources saved
GA4 as the panel“ChatGPT traffic is down”Optimizing for clicks you never measured in the answerPanel first, referrers lag

Crawl still matters on the retrieval path. If ChatGPT Search, Perplexity, and Gemini all go dark the week a WAF started challenging AI user-agents, start with fetch, not with new outlines. That eligibility work is documented on the playbook’s crawl spokes. This panel will show the outage. It will not explain robots.txt by itself.

The expensive failure is a sparkline with no owner. A KPI that cannot become a ticket is decoration.

How do I turn a four-engine miss into a ticket?

A miss that does not name a URL or a corroboration target will die in the deck. Map the hole to one object.

Engine × column holeLikely objectNot the ticket
ChatGPT mention no, Search onThird-party roundups and entity packetA new homepage slogan
ChatGPT recommend yes, owned cite noComparison / pricing URL the Sources panel can liftMore brand-name blog posts
Claude no-search mention wrongAbout + every profile that still has the old yearA Brave ranking campaign you cannot audit
Claude search on, citation noQuoteable primary page + fetch for Claude’s search/user agentsTreating n/a weeks as this miss
Gemini app mention noExtractable answer in HTML, not a client-only shellPasting GSC Overview impressions here
Perplexity mention yes, citation noUnique numbers and a table on a live URLBuying a Pages experiment
Perplexity citation yes, recommend noPick language and a shortlist table a numbered cite can liftChanging the prompt until you are first
All three retrieval engines dark the same weekrobots.txt / WAF / 403sA content calendar

Ticket hygiene:

  1. One miss → one owner → one URL or one corroboration ask.
  2. Re-run only the mapped prompt IDs the week the URL ships.
  3. If the cell does not move after the ship, the post did not change inclusion. It changed the CMS.

Do not open a ticket for a Claude no-search citation: n/a. That row is not a miss.

What should I skip if I only have a week?

Skip SOV math, skip a 40-prompt fantasy, skip buying a tracker before you have columns, skip blending Google Search into this sheet.

Do in week oneSkip
12–16 hire / best-of / alternative promptsBranded vanity prompts in the recommend rate
Four engines, one weekday, artifactsCopilot, AI Mode, Overviews in the same average
Mention / citation / recommend / accuracySentiment scores you will not act on
One competitor set sales acceptsIndustry “benchmark” citation %
Map each miss to a URL or a corroboration targetA content calendar of 20 outlines

Order for a short week:

  1. Write the scoring guide in the sheet header. Three columns, search_invoked, n/a rules for Claude.
  2. Lock 12 strings. Print them. Do not “improve” them on Friday.
  3. Run all four engines on the same day. 48 rows. That is a baseline, not a program.
  4. Pick three tickets from the worst engine × column pair. Ship those, not a blog about measurement.
  5. Put next Tuesday on a calendar with a named owner.

If you cannot finish 12 × 4, cut prompts, not engines. Four engines on eight prompts beats two engines on thirty you will abandon.

A week cannot prove a trend. It can prove you have a log. If week two does not exist on the calendar before week one starts, you are collecting souvenirs.

How do I report four engines without averaging them?

Bring a table, not a pie. The pie is how a Perplexity citation week hides a ChatGPT recommend collapse.

Slide object

  1. Panel v1, dates, market, frozen competitors.
  2. Counts first: n of N money prompts, per engine, for mention, citation, and recommend.
  3. The hole: engine × column that is actually losing.
  4. Three tickets with URLs.
  5. Ask: keep the weekly freeze, or book a visibility audit if nobody owns the log.
Exec questionBad answerBetter answer
“What’s our AI visibility?”“34%”“ChatGPT Search: 6/16 recommends, 3/16 owned cites. Claude: 4/16 mentions, 9/16 no-search. Gemini app: 2/16 mentions. Perplexity: 8/16 cites, 3/16 recommends. Panel v1, week of [date].”
“Are we healthy?”A vendor band“No portable healthy %. Here is three weeks on the same IDs.”
“Which engine first?”The one with a LinkedIn screenshotDeal-influence × current hole
“Did the rewrite work?”SessionsSame IDs, same engines, two cycles later
“Can we use Semrush’s number?”Yes, silentlyYes, and the prompt list, or we manage to the sheet

If a vendor number goes in the appendix, quote that product’s formula and engine coverage. Semrush Brand Performance is not this DIY event-count. Mixing them across quarters is how a “decline” appears that is really a formula change.

Do not put GSC generative AI impressions on this slide as if they were ChatGPT. Point to the measurement spoke for Google Search. Keep this slide for the four chat/answer UIs.

Four small multiples beat one average. Always.

When is a four-engine panel not worth running yet?

Skip it when the inputs are fiction, when buyers never use these products, or when nobody will sample twice. A weekly freeze with no second week is a workshop.

Skip this panel whenDo this firstThen start
You cannot name 12 buyer promptsA one-hour sales interviewLock panel_v1
Money URLs do not fetchCrawl, WAF, noindex accidentsAfter five URLs return in view-source
Legal blocks every fetcherA written allow-list for citation botsRetrieval engines only, labeled
You are mid-rebrandFreeze the legal name on AboutAfter entity pages agree
Buyers only use Google SearchRank + GSC gen-AI + Overview logAdd these four if a deal cites ChatGPT
Leadership wants one number and will not take fourRefuse the averageAudit, or no program
No owner for next TuesdayAssign or stopWeek 1 and week 3 on the calendar

DIY this if one named person will finish the weekly subset. Hire when the four engines disagree, the argument is political, and you need an outside prioritization hammer. Lane work is /visibility. The system map is the AEO playbook.

A visibility audit is the right next step when you cannot produce mention vs citation vs recommend, split by engine, after two honest weeks — not when you dislike the first percentages.

Do not wait for a vendor to invent a score you could have written down on Tuesday.

FAQ

How do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity?

Run the same frozen buyer prompts on each product every week and log mention, citation, and recommend as separate columns, plus whether search or grounding actually ran. Score each engine on its own row. There is no official four-engine console and no transferable “healthy” citation rate — your trend on frozen IDs is the measurement.

How do I measure whether do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity is working?

It is working when next week’s sheet uses the same prompt IDs, the same scoring guide, and artifacts you can reopen — and when a moved cell becomes a ticket on a URL. If the only output is a blended percentage, or if Claude no-search rows are scored as citation losses, the ritual is not working. Two cycles minimum before you brief a lift.

What usually fails first when teams try this?

They collapse mention, citation, and recommend into one score, then average the four engines. Next they mix ChatGPT Search-off with Search-on, and they treat a Claude answer that never searched as a Perplexity-style miss. Fix the scoreboard before you buy another tracker. Crawl failures still matter, but they show up as retrieval engines going dark together, not as a mysterious composite dip.

How long does this take to show results?

A baseline is one working day for 12–16 prompts across four engines. A trend takes at least two weekly cycles on frozen IDs — longer if you only sample monthly. Fetch and entity-packet fixes can move retrieval engines within weeks; they are not a citation SLA. Do not declare a win from one Tuesday.

What should I skip if I only have a week?

Skip share-of-voice math, branded vanity prompts inside the recommend rate, and any blend with AI Overviews or Copilot. Lock 12 strings, run all four engines in one day, save artifacts, and pick three tickets from the worst engine × column pair. Cut prompts before you cut engines if time runs out.

When is this not worth doing yet?

When you cannot name money prompts, when public URLs do not fetch, when buyers never use these four products, or when nobody will run the panel twice. Also skip it if leadership will only accept one blended number — that request is the opposite of this scoreboard. Keep Google measurement on its own spoke until a deal actually cites ChatGPT, Claude, Gemini, or Perplexity.

CTA

If you cannot see mention, citation, and recommend on ChatGPT, Claude, Gemini, and Perplexity as four logs, you cannot manage them. Freeze the panel. Stop averaging engines.

Lane: /visibility · Next step: visibility audit

FAQ

What questions does this article answer?

How do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity?
Run the same frozen buyer prompts on each product every week and log mention, citation, and recommend as separate columns, plus whether search or grounding actually ran. Score each engine on its own row. There is no official four-engine console and no transferable “healthy” citation rate — your trend on frozen IDs is the measurement.
How do I measure whether do I measure brand visibility in ChatGPT, Claude, Gemini, and Perplexity is working?
It is working when next week’s sheet uses the same prompt IDs, the same scoring guide, and artifacts you can reopen — and when a moved cell becomes a ticket on a URL. If the only output is a blended percentage, or if Claude no-search rows are scored as citation losses, the ritual is not working. Two cycles minimum before you brief a lift.
What usually fails first when teams try this?
They collapse mention, citation, and recommend into one score, then average the four engines. Next they mix ChatGPT Search-off with Search-on, and they treat a Claude answer that never searched as a Perplexity-style miss. Fix the scoreboard before you buy another tracker. Crawl failures still matter, but they show up as retrieval engines going dark together, not as a mysterious composite dip.
How long does this take to show results?
A baseline is one working day for 12–16 prompts across four engines. A trend takes at least two weekly cycles on frozen IDs — longer if you only sample monthly. Fetch and entity-packet fixes can move retrieval engines within weeks; they are not a citation SLA. Do not declare a win from one Tuesday.
What should I skip if I only have a week?
Skip share-of-voice math, branded vanity prompts inside the recommend rate, and any blend with AI Overviews or Copilot. Lock 12 strings, run all four engines in one day, save artifacts, and pick three tickets from the worst engine × column pair. Cut prompts before you cut engines if time runs out.
When is this not worth doing yet?
When you cannot name money prompts, when public URLs do not fetch, when buyers never use these four products, or when nobody will run the panel twice. Also skip it if leadership will only accept one blended number — that request is the opposite of this scoreboard. Keep Google measurement on its own spoke until a deal actually cites ChatGPT, Claude, Gemini, or Perplexity.
Sources

Last reviewed — OpenAI ChatGPT Search source UI, Anthropic web-search citations, Google AI-features docs, Perplexity Help Center citations, Semrush mention/cite/recommend definition, and Josh Blyskal 2026 Claude search-invocation study checked 2026-09-05.

More from this lane

AI Visibility

All →
Book the audit