Original Numbers Get Cited Because Models Hate Sharing Ambiguous Credit
Original research helps AI citations when numbers are dated, attributed, and extractable — GEO showed statistics and quotes lift answer visibility up to ~40%.
William Spurlock Founder — Spurlock Studios Updated 22 MIN
Yes — original research helps you get cited by AI answer engines when the number is yours, dated, method-labeled, and easy to lift in 40–80 words. Models prefer a clean statistic with a named source over three blogs restating the same vibes. Ambiguous credit loses.
This is a method spoke under the Answer Engine Optimization playbook. I have been SEO certified since 2021. The AEO version of that work is still one fact, one source, one URL.
The short answer
- Original, attributed statistics are among the strongest citeable assets you can ship
- Aggarwal et al.’s GEO paper (arXiv Nov 2023; KDD 2024) found evidence-adding edits — statistics, quotations, cite-sources — lifted visibility ~30–40% on their Position-Adjusted Word Count metric vs an unoptimized baseline
- That metric measures share of the generated answer, not traffic or leads. Do not sell it as “+40% organic”
- SMB-scale research counts: a 40-respondent customer survey beats another unattributed listicle
- Publish the number next to methods and date, then seed it through PR and answer-first pages
What the GEO study actually found
Cite the paper by name: Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, and Deshpande — “GEO: Generative Engine Optimization” (arXiv:2311.09735, Nov 2023; accepted KDD 2024, DOI 10.1145/3637528.3671900). Affiliations include Princeton and IIT Delhi (plus independent researchers).
What they measured, from the HTML preprint:
| Finding | What the paper reports | How to use it |
|---|---|---|
| Evidence-adding methods | Cite Sources, Quotation Addition, and Statistics Addition achieved ~30–40% relative improvement on Position-Adjusted Word Count vs baseline | Put real numbers, quotes, and sources on the page |
| Best-method ceiling | Best methods improved the baseline by 41% on PAWC and 28% on Subjective Impression in the main GEO-bench table | Treat “up to 40%” as a per-method upper bound, not your forecast |
| Live-engine check | On a commercial generative engine, Quotation Addition posted ~22% PAWC lift; Statistics Addition / Cite Sources showed lifts up to ~9% and ~37% across the two metrics | Statistics help live engines, not only the lab setup |
| Keyword stuffing | Offered little to no improvement in the main bench; ran ~10% worse than baseline on the live-engine slice | Stop stuffing; start attributing |
| Domain variation | Efficacy varies by domain (citations for factual queries; statistics for law and opinion) | Test your category; do not assume uniform lifts |
Hedge hard: these are relative visibility lifts inside their GEO-bench setup (≈10,000 queries, top-5 Google result text as sources) and a smaller live-engine slice (200 test samples, sources provided as files). Engines have moved since 2023–2024 models. Treat the direction as durable; treat the exact percentage as historical, not a guarantee for your domain in 2026.
How does original data actually get cited?
A model does not “award” you a citation because you announced a study. It retrieves pages, lifts a passage it can defend, and attaches a source when the product shows citations. OpenAI’s ChatGPT Search help is the operator check: inline citations and a Sources panel appear when search ran. No Sources control means the answer came from memory, not a live page. You cannot “research” your way into a training-weight hallucination the same way you research your way onto a retrieved URL.
Google’s AI features guidance is equally blunt. There is no extra markup for AI Overviews or AI Mode. A page must be indexed and snippet-eligible. Important content has to exist as text. Structured data must match the visible page.
The citation path for a number looks like this:
| Step | What has to be true | What fails the step |
|---|---|---|
| 1. Index | The research URL is crawlable, linked, and eligible for a snippet | Gated PDF only, noindex, or orphan URL |
| 2. Retrieve | The query is close enough that a retriever pulls your page | The number lives on a brand-essay with no question shape |
| 3. Extract | A 40–80 word block contains result + n + population + date | The stat is a chart with no alt text or a slide 14 |
| 4. Attribute | Only one owner can claim that exact measurement | Five agencies restated “most marketers struggle” |
| 5. Corroborate | A third-party page repeats the number and links you | You are the only domain that ever said it |
| 6. Cite | The product shows a supporting link or Sources chip | The model used the fact and dropped the URL |
Pew’s June 2026 Americans and AI study (surveyed Feb 17–23, 2026; 5,119 U.S. adults) found about half of U.S. adults use AI chatbots, 44% use ChatGPT, and six in ten read AI summaries at the top of search results. If those answers need a number, they will take the one they can attribute. If they cannot attribute it, they will average the vibes and cite nobody — or cite the highest-authority restatement.
- Research URL returns 200, is internally linked, and is not
noindex - Headline number is HTML text, not only an image
- Methods sit on the same screen as the number
- At least one cluster page links the canonical URL instead of restating a rounded figure
- At least one third-party mention exists or is in pitch
- Frozen prompt panel logged before and after publish
Why do models prefer a number they can attribute?
Credit assignment is the job. When five pages say “AI citations convert better” with no n, no date, and no method, a synthesizer has no reason to pick you. When one page says “In our March 2026 survey of 87 B2B SaaS marketers (methods below), 41% said…”, the claim is unique. Repeating it without the source looks like theft. Repeating it with the source looks like journalism.
Ahrefs’ 75,000-brand analysis put branded web mentions at a 0.664 Spearman correlation with AI Overview brand visibility, against 0.218 for referring domains. That is a correlation, not a causal law. It still explains why a number that other people repeat — with your name on it — outruns a homepage backlink with no sentence attached.
| Claim shape | Credit is | Typical model behavior |
|---|---|---|
| “Most teams struggle with AI search” | Shared / none | Restate, cite a roundup or nobody |
| Vendor “industry average” with no dataset | Contested | Cite the vendor, or drop the number |
| Your dated survey with n and methods | Yours | Cite the research URL if retrieved |
| Your number repeated on a trade site | Yours + corroborator | Cite either URL; both teach the fact |
| Invented n dressed as a national census | Toxic | Spreads, then poisons trust when checked |
Google’s helpful-content questions still ask whether the page provides original information, reporting, research, or analysis. Their AI-optimization guide says not to recycle what others already said, or what a model could invent. A unique number is the shortest version of that instruction.
Does original data improve AI citations?
Yes, when the data is hard to reassign. If five agencies all say “most marketers struggle with AI search” with no methods, the model has no reason to credit you. If you publish a measurement only you ran, you created a unique extractable claim.
Checklist for a citeable number:
- Sample size stated
- Population defined (who was asked or measured)
- Collection date or window stated
- Method in one paragraph (survey / log sample / scrape / cohort)
- Limitation stated (what it does not prove)
- Number appears in the first answer block, not only in a PDF
AAPOR’s disclosure standards exist for the same reason, even if you are not a polling firm: sample size, how the sample was generated, and dates of data collection belong next to the result. Hide those and you are asking a model to invent the provenance.
| Disclosure | Put it here | Why a model needs it |
|---|---|---|
| n | Lead sentence and table stub | Distinguishes a census from a hallway poll |
| Population | Same sentence as n | Stops the number from being applied to “everyone” |
| Window | Lead + temporalCoverage if you mark Dataset | Stale years become current “truth” |
| Method | Paragraph under the table | Survey ≠ scrape ≠ support-ticket count |
| Limit | One sentence after the headline | “Customer sample, not a national probability survey” |
What counts as “original” at SMB scale?
You do not need a 10,000-query academic bench. You need a number nobody else owns.
| Asset | SMB-feasible? | Citeability |
|---|---|---|
| Customer survey (n≥30 with honesty about limits) | Yes | High if methods sit next to the number |
| Internal ops benchmark (anonymized) | Yes | High for niche B2B |
| Price / feature matrix you maintain quarterly | Yes | High for comparison prompts |
| Support-ticket taxonomy with counts | Yes | High for “what actually breaks” prompts |
| Frozen 25-prompt answer-panel snapshot | Yes | High for category visibility claims |
| Scraped industry leaderboard (disclosed method) | Sometimes | Medium — disclose ethics and date |
| Fabricated “studies” | Never | Contaminates trust permanently |
A survey of your customers counts if you say so. “n=42 of our customers” is honest. Pretending it is a national census is fraud.
Google’s AI-optimization guide contrasts commodity listicles with first-hand, non-commodity pages. A 40-row customer survey is non-commodity. “7 tips for first-time AI search teams” is not, unless the tips are your measured failure modes.
How do you publish a number a model can extract?
Structure the page like a citation wants to be born:
- Lead with the number in the first 2–4 sentences
- Put methods and date on the same screen — not a separate PDF only
- Use a table for multi-stat findings
- Repeat the headline stat once in an FAQ H3 so FAQ extractors can grab it
- Link the canonical research URL from related how-tos instead of restating approximate numbers elsewhere
| Bad extract | Good extract |
|---|---|
| “Many teams see big gains from research.” | “In our April 2026 survey of 64 agency owners, 29 (45%) said AI answers already influenced at least one closed deal.” |
| Stats buried in slide 14 of a gated deck | Stats HTML-public, gated deep-dive optional |
| Undated “industry average” | Dated, attributed, limited |
| Chart image, no caption | Table + caption that restates n, window, and population |
| Rounded “about half” on ten cluster pages | Exact figure on the canonical URL; clusters link back |
Answer engines reward the second column. Humans do too. OpenAI’s ChatGPT Search product note is the user-facing version: answers come with links to news and posts so a person can click through. If your “research” is a vibe paragraph, there is nothing to click and nothing to quote.
Shape examples only — not Spurlock client results:
| Pattern | Use | Do not use |
|---|---|---|
| “In [window], among [n] [population], [result].” | Lead and table stub | A percentage with no denominator |
| “Methods: [mode], [recruit], [dates]. Limit: [what it is not].” | Same screen | A footer that says “details on request” |
| “Source: [canonical URL].” | Cluster spokes and pitches | Recopied rounded figures |
What belongs in the methods block?
Write the methods paragraph before you write the headline. If you cannot fill the table, you do not have a study. You have a slogan.
| Field | Minimum | Fail |
|---|---|---|
| Who | Named population (“our customers,” “US agency owners who replied”) | “Marketers” |
| How many | Completed responses or measured units | “Hundreds” |
| When | Start and end dates | “Recently” |
| How | Survey / logs / scrape / cohort, plus recruit path | “We asked around” |
| Weighting | None, or what you weighted and why | A margin of error on an opt-in sample with no model |
| Limit | One sentence | Silence |
| AI-generated rows | Explicit if any synthetic responses exist | Mixing bots into n |
AAPOR’s disclosure list also wants the method used to generate the sample and, for non-probability work, a warning against fake precision. An opt-in customer list is fine. A “±3%” on that list, with no model, is theater.
- Methods paragraph drafted before outreach
- n is completed units, not invites sent
- Opt-in / customer / convenience sample is labeled as such
- No margin of error unless you can defend the model
- Any AI-generated “respondents” are excluded or disclosed as non-human
- Anonymization pass done before the table goes public
I have shipped 500+ automations. The ops version of this block is the same as a runbook: who, when, what you counted, what you did not.
Does Dataset markup help citations?
Markup is not a citation cheat code. Google’s AI features page says there are no extra technical requirements beyond being indexed and snippet-eligible. Dataset markup can still help discovery of the dataset and keep machines from inventing your provenance — if it matches the visible text.
Google’s Dataset structured-data docs and schema.org Dataset are the pair to use when the page is actually a dataset landing page, not a how-to with one stat buried in paragraph four.
| Property | Put on the page | Markup note |
|---|---|---|
name | H1 or dataset title | Required for Dataset rich results |
description | 50–5,000 characters that include n, window, and limit | Do not write a slogan |
creator / publisher | Your legal or common brand name | Must match the byline |
temporalCoverage | Collection window | Same dates as the lead |
measurementTechnique | Survey, log extract, scrape | Same words as methods |
variableMeasured | What you counted | One property per headline metric |
identifier | Canonical URL or DOI if you have one | Do not mint a fake DOI |
citation | Related papers you want cited in addition | Do not use this field to cite yourself |
Google’s Dataset docs are explicit: do not use citation for the dataset’s own cite. Use name, identifier, creator, and publisher for that. If the JSON-LD says n=200 and the table says n=40, you trained the conflict yourself.
- Visible table and JSON-LD agree on n, dates, and title
- Dataset markup only on a page that is actually a dataset
- Article / BlogPosting still describes the write-up
- No invented DOI
-
isAccessibleForFreeis honest if you gate the CSV
How do you seed a number without a PR team?
A unique number that never leaves your domain is still better than a fake one. It is weaker than a unique number other people repeat. Ahrefs’ mention correlation is the reason: the web has to say your name next to the fact.
- Publish the canonical page with schema-honest Article (and Dataset, if it qualifies) markup
- Pitch three niche newsletters that already cover your category — one sentence + the number + the URL
- Reply in relevant Reddit / community threads only where the number answers the question asked
- Send the page to partners who cite stats in their own posts
- Refresh the number on a calendar, not when you feel anxious
For journalist-shaped amplification, use PR and digital PR for citations. That spoke is the corroboration layer. This one is the atom they pitch.
| Channel | What you send | Success looks like |
|---|---|---|
| Niche newsletter | One sentence, n, date, URL | They quote the number and link you |
| Trade explainer | Table + methods paragraph | A criteria row sourced to your URL |
| Partner blog | Canonical link, not a rewritten approx | One source of truth |
| Community thread | The number only if it answers the asked question | No drive-by study dumps |
| Your cluster | Internal link, exact figure on the research URL | Spokes do not drift to “about half” |
- Pitch subject line contains the number, not “thoughts on AI search”
- Editor packet includes methods and the limit sentence
- You are willing to see the quote live forever
- Paid vs earned is labeled if anyone paid for the placement
- Correction contact is a same-day inbox, not info@
What breaks when the statistic is unsourced?
What breaks: a blog claims “AI citations convert 4.4× better” with no study link. Competitors copy it. Models repeat it. Your brand becomes the rumor’s origin — or worse, gets none of the credit while the fake number spreads.
What it costs: credibility with operators who check sources, and hallucination cleanup later. Once a fake n is on three domains, corrections lag syndication.
What you do instead: publish only numbers you can defend, or hedge explicitly (“vendor claim; we have not verified”). Cut the rest.
| Failure | How it starts | Cleanup |
|---|---|---|
| Orphan rumor | You tweeted a number with no URL | Publish the methods page or retract |
| Stolen credit | A roundup restated you without a link | Ask for a correction; keep the canonical live |
| Hallucinated n | A model rounded 41% to “most” and dropped you | Re-run the panel; strengthen the lead block |
| Contagion | Five blogs copied an unsourced “4.4×” | Do not add a sixth. Publish a real measurement or stay quiet |
| Trust kill | You invented n and got caught | Retract. Do not “update the graphic” |
The FTC’s August 2024 final rule on fake reviews and testimonials (effective October 21, 2024) is about reviews, not AEO. The direction still applies: AI-generated fake social proof is a consumer-protection problem, not a content tactic. Invented survey respondents sit in the same family.
What should you never invent?
Never publish:
- Invented sample sizes
- “Industry averages” with no dataset
- Competitor revenue guesses presented as measurement
- AI-generated survey respondents counted as humans
- Recycled vendor claims rebranded as your study
- A margin of error on a convenience sample with no model
- A national-census frame around a customer list
If legal would not put the number in a pitch deck for a serious buyer, do not put it on a citation page. Hallucinated research is worse than thin content — it teaches models the wrong fact with your URL attached.
| Temptation | Honest substitute |
|---|---|
| “Thousands of marketers say” | “n=40 of our customers, June 2026” |
| “Industry average close rate” | “Median days-to-first-live across our last 12 projects (anonymized)” |
| “4.4× citation conversion” | Cut it unless you have the study URL |
| Synthetic panel dressed as humans | Disclose non-human rows, or delete them |
| Competitor’s ARR “estimate” | Feature/price matrix you actually collected |
AAPOR’s 2026 transparency materials now ask researchers to say whether a sample was AI-generated. If you cannot say “all rows are human,” you do not have a customer survey.
How often should you refresh a benchmark?
| Cadence | Fits |
|---|---|
| Quarterly | Fast-moving tooling / pricing markets |
| Semi-annual | B2B process benchmarks |
| Annual | Large surveys that are expensive to rerun |
| Event-driven | After a platform shock (major AI Overview change, new engine) |
Stale dates kill trust. A 2023 survey presented as current truth is worse than no survey. Put the year in the H1 or lead.
| Refresh trigger | Action |
|---|---|
| Window is older than the cadence | Rerun or stamp “historical, collected [dates]” |
| Population changed (new ICP, new offer) | New study, new URL or new version field |
| Method changed | Do not splice old n into new n without saying so |
| Engine behavior shifted on your panel | Re-baseline citations; do not silently edit the old table |
| You cannot defend a cell | Blank it. Do not interpolate |
- Collection window in the H1 or first paragraph
-
lastModified/dateModifiedmatches a real edit - Old PDFs stamped superseded
- Cluster pages still point at the current canonical
- Panel re-run scheduled on the same calendar as the refresh
How do research, clusters, and PR share one number?
Original research is the atom. Clusters distribute it. PR corroborates it.
| Layer | Job |
|---|---|
| Research page | Own the number |
| Cluster spokes | Apply the number to buyer questions |
| Digital PR | Get third parties to cite your URL |
| Measurement | Log when answers cite your research URL |
Do not build fifteen posts that each invent a new fake statistic. Build one honest dataset and cite it everywhere.
| Page type | Role of the number |
|---|---|
| Research canonical | Full methods + tables |
| Definition spoke | One lined statistic in the lead |
| Comparison spoke | Table cell sourced to research URL |
| How-to spoke | “We measured X; therefore step 2…” |
| PR placement | The same sentence, not a rounded cousin |
Duplicate the headline number sparingly. Prefer linking back to the canonical research URL so models learn one source of truth. If a spoke says “about half” and the study says 41%, you created a conflict the synthesizer will average away — and then it has no one to cite.
A 14-day SMB research recipe
Day 1–2: pick one question buyers ask that has no owned number. Day 3–5: run a short survey or pull an anonymized internal sample (n you can stand behind). Day 6–8: write the research page with lead number, methods, table, limitations. Day 9–11: update two existing posts to cite the new page. Day 12–14: pitch three outlets / newsletters; re-run your AI prompt panel and log citations.
- One research question locked
- Methods paragraph written before outreach
- Canonical URL live
- Two internal links from cluster pages
- Panel re-baseline archived
| Day | Output | Done means |
|---|---|---|
| 1–2 | One question + population | You can write the lead sentence |
| 3–5 | Raw table, anonymized | n is completed units |
| 6–8 | Public HTML page | Number + methods on one screen |
| 9–11 | Two cluster edits | Exact figure or a link, not a new approx |
| 12–14 | Three pitches + panel log | URLs and quoted sentences stored |
Ship the number. Then argue about the number. Models follow the argument that has a receipt.
Cheap research ideas that still count
- Support-ticket taxonomy — top 10 reasons customers contact you this quarter (anonymized counts)
- Time-to-X benchmark — median days from kickoff to first automation live across your last N projects (anonymized)
- Feature usage cut — % of accounts using the one feature buyers ask about
- Price-of-inaction diary — hours clients logged before vs after a workflow (opt-in, aggregated)
- Answer-panel snapshot — how often competitors appear in AI answers for a fixed 25-prompt set (your measurement, dated)
Each of these is original if you collected it. None require a university IRB. All require a methods paragraph.
| Idea | Population | Typical limit to print |
|---|---|---|
| Ticket taxonomy | Your inbox, one quarter | Not “the industry’s top issues” |
| Time-to-X | Your last N projects | Selection bias; you choose clients |
| Feature cut | Accounts you can query | Product-specific, not category-wide |
| Hours diary | Opt-in clients only | Survivorship; quiet failures drop out |
| Prompt panel | Your 25–40 frozen prompts | Engine + date + prompt text required |
I have 20,000+ hours in agentic systems and 35,000+ hours saved for clients across the book of work. Those are receipts I will put next to a method. They are not a substitute for a dated table about your category.
Quote + statistic pairing (what GEO rewarded)
The GEO methods that worked were not “more keywords.” They were evidence: statistics, quotations, and source citations. Practically:
| Element | On-page pattern |
|---|---|
| Statistic | “In [window], among [n] [population], [result].” |
| Quotation | Named expert or customer quote next to the claim it supports |
| Cite sources | Outbound links to primary data you did not invent |
You can ship all three on one page without turning into an academic journal. Keep the lead human; keep the receipts dense.
| Pairing | Why it extracts |
|---|---|
| Number + named customer quote | Statistic for the table; quote for the “what it felt like” sentence |
| Number + outbound primary source | You are not pretending you ran Pew’s survey |
| Number + your limit sentence | Stops the model from promoting you to a census |
| Quote with no number | Fine for color; weak as the only asset |
| Number with no source | Fine if it is yours; label it |
GEO’s keyword-stuffing arm is the control you should remember: extra query words did not buy answer share. Evidence did.
How do you measure whether the number got cited?
If you do not log citations, you are collecting vibes. Freeze 25–40 buyer prompts before you publish. Run the same panel 4–6 weeks later. Record the engine, date, whether search/citations appeared, and whether your research URL or brand was named.
| Log column | Why it exists |
|---|---|
| Prompt (verbatim) | The panel must not drift |
| Engine + product surface | ChatGPT Search ≠ AI Overview ≠ AI Mode |
| Search/citations visible? | Memory answers are a different problem |
| Your research URL cited? | Primary win |
| Brand mentioned, URL missing? | Partial; pitch corroboration |
| Competitor URL cited? | Gap, not a vibe |
| Quoted sentence | What actually got lifted |
OpenAI’s help page is the ChatGPT check: hover or open Sources. Google Search Console’s generative-AI reports (where you have them) are the Search check. Neither replaces a dated spreadsheet you own.
- Panel frozen before publish
- Baseline archived with screenshots or exports
- Re-run scheduled 4–6 weeks out
- Wins are URLs, not “we felt more visible”
- Losses become tickets (extractability, corroboration, or index)
Inclusion is the KPI. Domain Rating on the research subdomain is a different product.
FAQ
What did the GEO study actually find about statistics?
Aggarwal et al. (GEO, arXiv Nov 2023 / KDD 2024) reported that evidence-adding methods — including Statistics Addition, Quotation Addition, and Cite Sources — produced roughly 30–40% relative gains on Position-Adjusted Word Count versus an unoptimized baseline. On their live-engine slice, Quotation Addition posted about 22% PAWC lift, and Statistics Addition / Cite Sources showed lifts up to about 9% and 37% across the two metrics. That is answer-visibility share, not traffic. Keyword stuffing underperformed the baseline in their tests.
Do I need a huge sample size?
No. You need honesty. A clear n=40 customer survey with limitations beats a vague “thousands of marketers say” claim. State who was surveyed and what you cannot conclude. AAPOR’s disclosure standards still want sample size and recruit method next to the result, even on a small convenience sample.
Should methods and dates sit next to the number?
Yes. Put sample, population, date, and method on the same screen as the headline statistic. Buried methods get stripped; undated stats age into lies. If you add Dataset markup, those fields must match the visible table.
How often should I refresh a benchmark?
Match the market’s rate of change — often quarterly for tooling, semi-annually or annually for slower B2B process data. Always show the collection window in the lead. A 2023 survey presented as current truth is worse than no survey.
Can a survey of my customers count?
Yes, if you label it as your customer sample. That is original. Pretending a customer sample is a national probability survey is not. Do not invent respondents, and do not count AI-generated rows as humans.
How does research interact with digital PR?
Research gives PR a sentence worth pitching. PR gives the research third-party URLs that models can corroborate. Neither replaces the other — see digital PR for citations. Ahrefs’ mention-vs-link split is why the third-party sentence matters.
CTA
Own a number nobody else can claim — dated, method-labeled, and public.
Lane overview: /visibility. Next step: a visibility audit.
What questions does this article answer?
- What did the GEO study actually find about statistics?
- Aggarwal et al. (GEO, arXiv Nov 2023 / KDD 2024) reported that evidence-adding methods — including Statistics Addition, Quotation Addition, and Cite Sources — produced roughly 30–40% relative gains on Position-Adjusted Word Count versus an unoptimized baseline. On their live-engine slice, Quotation Addition posted about 22% PAWC lift, and Statistics Addition / Cite Sources showed lifts up to about 9% and 37% across the two metrics. That is answer-visibility share, not traffic. Keyword stuffing underperformed the baseline in their tests.
- Do I need a huge sample size?
- No. You need honesty. A clear n=40 customer survey with limitations beats a vague “thousands of marketers say” claim. State who was surveyed and what you cannot conclude. AAPOR’s disclosure standards still want sample size and recruit method next to the result, even on a small convenience sample.
- Should methods and dates sit next to the number?
- Yes. Put sample, population, date, and method on the same screen as the headline statistic. Buried methods get stripped; undated stats age into lies. If you add Dataset markup, those fields must match the visible table.
- How often should I refresh a benchmark?
- Match the market’s rate of change — often quarterly for tooling, semi-annually or annually for slower B2B process data. Always show the collection window in the lead. A 2023 survey presented as current truth is worse than no survey.
- Can a survey of my customers count?
- Yes, if you label it as your customer sample. That is original. Pretending a customer sample is a national probability survey is not. Do not invent respondents, and do not count AI-generated rows as humans.
- How does research interact with digital PR?
- Research gives PR a sentence worth pitching. PR gives the research third-party URLs that models can corroborate. Neither replaces the other — see [digital PR for citations](/blog/pr-and-digital-pr-for-citations). Ahrefs’ mention-vs-link split is why the third-party sentence matters.
Last reviewed — GEO arXiv:2311.09735 / KDD 2024; Google AI-features, helpful-content, AI-optimization, and Dataset markup; schema.org Dataset; Ahrefs 75k-brand mention study; Pew Americans and AI 2026; OpenAI ChatGPT Search help; AAPOR disclosure standards; FTC fake-review rule checked 2026-08-16.
AI Visibility
AI Visibility Cannabis visibility when the ad accounts are banned
Google and Meta will not take the usual spend. The models still answer dispensary, cultivator, and brand questions — if the site can be read and the cart can clear a 21+ order.
AI Visibility When ChatGPT names the franchise, not your shop
Run the best-HVAC-near-me prompt panel. If the model names a national franchise, fix corroboration and entity facts — not another blog calendar.
AI Visibility How do I get cited by Perplexity specifically
Allow PerplexityBot, put a liftable answer and unique numbers in HTML, then log numbered sources on a frozen prompt panel. There is no bought citation rate.
AI Visibility What belongs in an AI visibility monthly retainer vs a one-time audit
A one-time audit is the baseline plus prioritized fixes. A monthly retainer is prompt-panel tracking, entity hygiene, page jobs, and citation recovery.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.