A Screenshot Is Not a Measurement
Somewhere right now a customer is asking an assistant who to call, and the answer either has your name in it or it doesn't. In 2026, Forrester reported that in every country it surveyed, more than half of online adults have used generative AI as a new way to find answers. It's not only chatbots, either. Similarweb's data, published in 2026, puts AI Overviews on more than 40 percent of US searches, up from about a third a year earlier, and counts 9.5 billion monthly visits to generative AI platforms as of May 2026, a 70 percent jump in a year.
So most owners do the sensible thing. They open an assistant, ask for a business like theirs, and screenshot the answer. That's a fine first look, and our guide to checking whether AI recommends your business walks through it. But it isn't tracking, and here's why. AI answers don't hold still. In a 2026 study of nearly 700,000 answers, Parse found that two ChatGPT answers to the identical prompt a month apart shared only 21.2 percent of their cited sources. Google AI Overviews were steadier and still shared only 31.5 percent. Shrinking the gap to a week barely helped.
The same research also shows the way out. When BrightEdge tracked thousands of prompts in February 2026, 96.8 percent of cited domains held their aggregate position week to week. Individual answers are noisy. The average across many answers is stable enough to trust. That's the whole trick to tracking AI visibility. You stop asking whether you were named today and start measuring how often you're named, across a fixed set of questions, on a schedule.
Four Numbers That Describe Your AI Visibility
Rank tracking gave you one number, your position. AI visibility needs four, because an assistant can name you without linking you, link you without naming you, and describe you wrong while doing either. Each metric is a simple rate over the same pool of answers.
| Metric | How to calculate it | What it tells you |
|---|---|---|
| Mention rate | Answers that name your business, divided by all answers you ran | Whether you're in the consideration set at all |
| Share of voice | Your mentions, divided by mentions of you plus each competitor you track | Who's winning the same questions, and by how much |
| Citation rate | Answers that link to a page on your site, divided by all answers | Whether your site is a source the assistant trusts |
| Accuracy | Answers that name you and get your facts right, divided by answers that name you | Whether the pitch a customer reads is the one you'd give |
Mention and citation have to stay separate columns, because they move independently. In 2026, Semrush's Ghost Citations study logged 3,981 brand appearances across 115 prompts and found that 61.7 percent were citations with no mention, the assistant used the site as a source and never said the brand's name. Only 13.2 percent were both cited and named, and 25.1 percent were named with no link. The platforms even lean opposite ways. Gemini named the brand in 83.7 percent of its appearances but linked the site in only 21.4 percent, while ChatGPT ran the other direction, citing in 87 percent and naming in about 21.

Accuracy is the one owners skip and shouldn't. If the assistant names you and says you don't do emergency calls when you do, that mention costs you the job. Track it as a plain yes or no per answer. And keep an eye on which pages get cited, since that reveals where AI gets its data about your business and which of your pages are doing the work.
Twelve Prompts Are Your Baseline
Tracking needs a fixed question set, the same prompts every time, or your numbers aren't comparable from one month to the next. Twelve is a workable number for a single-location business. Build them from the intents a real buyer has, not from how you describe yourself, and write three phrasings of each so a quirk of wording doesn't decide your score.
Take a worked example, Ridgeline Plumbing in Boise. Their customers arrive with four intents. An emergency, a replacement, a specific repair, and a comparison shop. Three phrasings each gives twelve prompts.
EMERGENCY
1. emergency plumber in Boise
2. who can fix a burst pipe tonight in Boise, Idaho
3. 24 hour plumber near me Boise
REPLACEMENT
4. best water heater installer in Boise
5. who should I hire to replace a water heater in Boise
6. tankless water heater installation Boise reviews
REPAIR
7. plumber for a slab leak in Boise
8. who fixes low water pressure in Boise homes
9. drain cleaning company Boise recommended
COMPARISON
10. top rated plumbers in Boise Idaho
11. Ridgeline Plumbing vs other Boise plumbers
12. most trusted plumbing company in BoiseNotice what's in the set. Real neighborhoods and problems, the words people say out loud, and exactly one prompt with the brand name in it. That branded prompt measures something different, whether the assistant knows you exist and what it says when asked directly. Keep it, but don't let it flatter your average. If you serve several towns, swap the city in three or four prompts rather than adding more, so the set stays small enough to run by hand.
Sample Enough Answers to Trust the Change
Here's the part most tracking advice glosses over. Because each answer is a coin flip weighted by your real visibility, a small sample can swing wildly without anything changing underneath. In 2026, Cloro worked the statistics. A visibility rate measured over 72 answers carries a margin of roughly plus or minus 10 percentage points. You need about 300 answers in a period to narrow that to about 5, and about 294 answers on each side to say with confidence that a rate moved from 20 percent to 30. One prompt checked daily gives you seven answers a week, which resolves nothing.
Now apply that to Ridgeline. Twelve prompts across each assistant we track, ChatGPT, Gemini, Claude, Perplexity, and Google AI, is 60 answers per pass. One pass a week is roughly 240 answers a month, close to the mark where a rate is worth reading. So a hand-run scorecard should compare month to month, never week to week. A move from 35 percent to 38 is noise. A move from 35 to 50 is real, and so is a slide from 35 to 20.

The cadence tradeoff is real. Weekly passes with a small set give you a monthly read. If you want to catch a slip within days, say after a competitor floods a directory with reviews, you need either far more prompts or daily sampling, which is past what anyone does by hand. That split is the design behind Lighthouse Local's AI visibility tracker, a weekly sweep across your whole prompt pool plus daily checks on the handful of prompts that matter most, so the monthly trend and the early warning come from the same data.
Keep the Scorecard in a Spreadsheet
The scorecard is one row per answer, and the four metrics fall out of it with a sum and a division. Nine columns are enough. Date, platform, prompt, named (1 or 0), position (first, middle, last, or none), competitors named, cited (1 or 0), facts correct (1 or 0), and a notes column for the wording that surprised you.
Ridgeline's first weekly pass looked like this. Named in 21 of 60 answers, a 35 percent mention rate. Across the same 60 answers, Treasure Valley Plumbing was named 33 times and Boise Drain Pros 18, so Ridgeline's share of voice against those two was 21 out of 72, about 29 percent. Their site was linked in 7 answers, a 12 percent citation rate. And of the 21 answers that named them, 17 had the facts right, an 81 percent accuracy score. The four wrong ones all said the same thing, that Ridgeline doesn't offer emergency service, which points straight at a page to fix.
You don't have to build any of this yourself. Paste the following into an assistant, filled in for your own business, and it will hand you the prompt set and the sheet.
I run [business type] called [business name] in [city, state]. I want to track how often AI assistants recommend us.
1. Write 12 prompts a real customer would type, 3 phrasings each for these intents: [emergency need], [big purchase or replacement], [specific repair or service], [comparison shopping]. Use the real city. Exactly one prompt should include our business name.
2. Build a spreadsheet template with these columns: Date, Platform, Prompt, Named (1/0), Position (first/middle/last/none), Competitors named, Cited (1/0), Facts correct (1/0), Notes.
3. Add formulas for Mention rate, Share of voice against [competitor 1] and [competitor 2], Citation rate, and Accuracy, calculated per month.
4. Tell me which of the 12 prompts to check daily if I only have time for three.Then run the twelve prompts through each assistant once a week, paste the answers into the sheet, and mark the four columns. It takes about an hour. The first pass is your baseline, and nothing before it counts.
Follow the Clicks in Analytics and Search Console
The scorecard measures the answer. Two free reports measure what happens after it, and both arrived in 2026. On May 13, 2026, Google Analytics added a dedicated AI Assistant channel. When a visitor arrives from a recognized assistant, the session's medium is set to ai-assistant automatically, and it lands in its own channel in your default channel group with no setup on your end. Google names a few assistants as examples and doesn't publish the full referrer list, so it's worth also filtering your referral report for any assistant you track that isn't showing up in the new channel.
On the Google side, Search Console announced a generative AI performance report on June 3, 2026, and rolled it out to every site worldwide by August 31. It breaks out your impressions inside AI Overviews and AI Mode, data that was already folded into your overall performance totals but never separated. Watch it monthly beside your scorecard. If the assistant is showing your page more and your mention rate is rising too, the two are confirming each other.
Keep the click numbers in proportion, though. AI referrals are growing fast, Adobe Analytics measured a 693.4 percent year-over-year jump in traffic to retail sites from generative AI tools over the 2025 holiday season, but for a local service business the count will still be small. Most of the value of being named is a phone call, and no analytics tool sees the customer who read your name in an answer and dialed. That's why the answer-side metrics stay primary, and why a plain habit of asking new customers how they found you belongs in the same monthly review.
By Hand or With a Tracker
Be honest about which camp you're in. The spreadsheet works well for one location, twelve prompts, and a monthly read, and if that describes you, run it for three months before you spend a dollar. It stops working in three predictable places. When you add locations and the prompt set triples. When you need to know about a slip within days rather than at month's end. And when you're an agency running this for many clients, where an hour a week per client turns into someone's whole job.
That's where a tool earns its cost. Lighthouse Local for Business runs the sweep for you across every assistant we track, tags every answer as mentioned, not mentioned, or competitor won, shows which of your pages each assistant cited, and benchmarks your mentions against each rival's on every prompt. The history is kept, so the trend line is there when you need it instead of rebuilt from screenshots. If you'd rather not do the fixing either, its specialists handle the on-page work, reviews, and listings that the wrong answers point at, with you approving each step. The Lighthouse Local for Business page lays out what's included at each level.
One distinction worth keeping straight. Our free AI visibility audit is a different tool from the tracker. It scans your site for the signals assistants read and returns a ranked fix list, which is the right first step when your accuracy score is low or your citation rate is near zero. It doesn't read what the assistants say about you. The scorecard, or the tracker, does that.
Watch the Line, Not the Snapshot
Tracking AI visibility comes down to a discipline, not a dashboard. Fix the questions, run them on a schedule across each assistant, score every answer on the same four columns, and only draw conclusions when the sample is big enough to carry them. Do that and the noise that makes screenshots useless becomes a trend line you can act on, the wrong facts become a to-do list, and the competitor winning your best prompt becomes a name you can study. The work to actually move the numbers is covered in our guide to getting your business recommended by AI. Whether you keep the sheet yourself or let our tracker keep it, the rule is the same. One answer tells you what an assistant said once. A hundred tell you where you stand.
Frequently asked questions
How do you measure brand visibility in AI search?
Run a fixed set of customer-style prompts through each assistant on a schedule, then score every answer on four rates. Mention rate, share of voice against competitors, citation rate, and accuracy. Because single answers vary a lot, compare the rates month to month over a few hundred answers, not one screenshot to the next.
What is a good AI mention rate?
There's no universal benchmark, and it depends on how many competitors share your market. Your own baseline is what matters. Measure it, then judge yourself against your best competitor's rate on the same prompts and against your own number three months later.
How often should you track AI visibility?
Run your prompt set weekly and read the results monthly. A dozen prompts across the assistants you track is about 60 answers a pass, so four passes get you near the 300 answers it takes to trust a change of a few points. Daily reads only make sense with a tool sampling far more answers.
Can Google Analytics show traffic from AI assistants?
Yes. Since May 2026, Google Analytics groups visits from recognized AI assistants into an AI Assistant channel automatically, with the medium set to ai-assistant. Search Console also has a generative AI report showing impressions in AI Overviews and AI Mode.
Do I need a tool to track AI visibility?
Not at first. A spreadsheet handles one location and a monthly read. A tool pays off when you add locations, need to catch changes within days, or run tracking for many clients, since that's when hand sampling can't reach a big enough sample.



