Table of Contents
- Key takeaways
- What are the four AI visibility metrics to track?
- How to build a prompt set for AI search visibility tracking
- Visibility: how often your client shows up in AI answers
- Position: where your client surfaces in an AI answer
- Sentiment: how AI engines describe your client
- Citations: which pages an AI answer was built from
- How to track AI search visibility in AgencyAnalytics
- Add AI visibility to your client reports
- Frequently asked questions about reporting AI search visibility to clients
7,000+ agencies have ditched manual reports. You can too.
Free 14-Day TrialTable of Contents
- Key takeaways
- What are the four AI visibility metrics to track?
- How to build a prompt set for AI search visibility tracking
- Visibility: how often your client shows up in AI answers
- Position: where your client surfaces in an AI answer
- Sentiment: how AI engines describe your client
- Citations: which pages an AI answer was built from
- How to track AI search visibility in AgencyAnalytics
- Add AI visibility to your client reports
- Frequently asked questions about reporting AI search visibility to clients
7,000+ agencies have ditched manual reports. You can too.
Free 14-Day TrialFor most marketing agencies, the question comes from a client right before a meeting: "Hey, by the way, are we showing up in ChatGPT?"
It's a fair thing to ask, and there's no good way to answer it on the spot. So you open a tab, run a few prompts, find one where the client's name lands in the output, and screenshot it for Friday's deck. Seems like a perfectly serviceable answer, for now.
The trouble shows up next month, when you do it again, and nobody can tell whether anything changed. The prompts were a little different. So was the phrasing, and the day of the week, and probably the account you were logged into. There's no line to draw between the two screenshots, because there was never a fixed starting point.
You’re not the only agency fielding that question. In our 2026 Marketing Agency Benchmarks Report, 66% of the 494 agency professionals we surveyed said clients are requesting AEO and SEO for AI search, making it the most requested new service of the year, ahead of performance-based paid ads and short-form video.

Most agencies are working out the answer as they go. That makes AI search visibility the new attribution problem for agencies: if you can't consistently measure presence in AI-generated answers, you can't prove progress, guide content strategy, or build trust in the work.
Key takeaways
Lock your prompt set before your first report. The fixed list of questions you run each month is the new keyword list. If it changes between reports, this month's number can't be compared to last month's.
Report visibility and position per engine, never blended. One averaged number hides which engines are gaining and which are slipping. Your client will eventually check one and find a number that doesn't match your report.
Always pair sentiment with the response that produced it. A client can be highly visible and still described with outdated information, so a score alone gives them nothing to act on.
Treat citations as the to-do list. Sort the sources feeding AI answers into owned, earned, and competitor pages, then flag an action beside each one.
AI Tracker in AgencyAnalytics keeps the inputs steady for you. It runs the same set of prompts across six AI engines and reports the results alongside your client's traffic, conversions, and campaign performance, giving them the full picture of performance.
What are the four AI visibility metrics to track?
To accurately report AI search visibility to clients, track four metrics per engine against a fixed set of prompts: visibility (how often the brand appears), position (where it lands in the answer), sentiment (how it's described), and citations (which sources the answer was built from). Report monthly, and show the results by engine.
Metric | What it measures | What it tells you |
Visibility | The share of tracked prompts where the brand appears in the answer. | Whether your client is in the conversation at all. |
Position | Where the brand lands in the answer: early, late, or not at all. | Whether buyers see your client’s brand before their competitor’s. |
Sentiment | How the AI engine describes the brand when it’s named. | Whether the description is accurate or current. |
Citations | The URLs the engine pulled from to build the answer. | What to write, refresh, or earn next. |
How to build a prompt set for AI search visibility tracking
Before you report anything, lock your prompt set. AI search has a lot in common with search engine optimization (SEO). Measuring it doesn't. Rank tracking gave you a stable thing to point at, and this doesn't work that way.
Most agencies have already run into this. Tracking people who discover a brand through AI tools is now the single biggest attribution challenge in the industry, named by 48% of the 494 agencies we surveyed in the 2026 Marketing Agency Benchmarks Report. Traditional analytics can't see most of it, which is why the measurement has to start with the answer itself rather than the traffic it sends.
Finding ways to report on how clients are showing up in AI results is going to be critical to our success.
Rachel Li, Marketing Director, Swifty
Think of your prompt set as the new keyword list. Everything you report comes out of it, and a sloppy set will hand you numbers that move for reasons that have nothing to do with the work you did.
There are five decisions worth making before your first report. Let’s use a hypothetical client as an example: a commercial roofing company in Austin, tracked across 30 prompts.
Choose prompts by buying stage, not by brand name
"Tell me about [client]" will almost always return the client because you've given the AI engine the answer. It looks great in a screenshot, but it tells you nothing.
The prompts worth tracking are the ones a buyer types before they know who to call:
"Best commercial roofing company in Austin"
"How much does commercial roof replacement cost?"
"TPO vs EPDM for a flat roof"
Most of your set should be unbranded questions about the problem, with a few branded ones to see how your client gets described when they do come up.
Group prompts into topics
Forty prompt rows on a client report is a wall of data nobody reads. Four topics with performance under each is a conversation.
Topics also surface patterns you'd miss otherwise. Say a client turns up for every pricing question and vanishes on comparison questions. That's a clear content gap you can fix.
Lock engine and location
Answers vary by engine and by where the question appears to be coming from. Track a local services client from the wrong city, and you'll get numbers that are technically accurate and pretty much useless.
Pick your engines, pick your location, and write both down somewhere you'll find them next month.
Freeze the set
Add prompts mid-quarter, and you've changed the total you're measuring against. A percentage built on a moving total can't be compared to last month's.
The set will need to change eventually, and that's fine. When it does, save the old one, date the new one, and note the change on the report so nobody has to reconstruct it later.
Report monthly, not weekly
Run the same prompt twice in an afternoon, and you may get two different answers. That's how generative AI works, and it means a single day's reading is mostly noise.
Sample over weeks and report monthly. This one decision will spare you from walking a client through a drop that was never real, which is not a conversation anyone wants to have twice.
The prompt set is the input. Visibility, position, sentiment, and citations are what you report out of it, in that order.
Visibility: how often your client shows up in AI answers
Visibility is the share of your tracked prompts where the client's brand appears in the response, measured per engine over a set period. It’s sometimes referred to as brand mention rate.
The math is simple enough. Count the responses where the brand shows up, divide by the total, and do it per engine rather than blended. The gap between engines is usually wide enough that a single average hides everything interesting.
Report how many prompts the percentage is based on, every time. 40% across 25 prompts and 40% across 200 prompts are not the same result, and the second one took far more work to earn. Early on, the direction matters more than the number itself, and it's worth saying so before the first report rather than after the second.
Here's the shape it should take:
Engine | Responses with a mention | Responses collected | Visibility | Change vs. last month (responses) |
ChatGPT | 18 | 30 | 60% | +6 |
Google AI Overviews | 12 | 30 | 40% | +4 |
Perplexity | 3 | 30 | 10% | -2 |
Blended | 33 | 90 | 37% | +8 |
If you read the bottom row on its own, this looks like a good month: 37%, up eight responses. The rows above it say something else. ChatGPT is carrying nearly all of the gain, and Perplexity is going the other way.
What to say in the meeting: "In Google AI Overviews, you're appearing in 12 of the 30 questions we track, up from eight last month, and most of that gain is on comparison questions."
Watch out for a single blended number across all engines. It hides where the work is, and creates confusion as soon as your client checks a single engine and sees something different.
Position: where your client surfaces in an AI answer
Position in an AI answer is where the brand appears in the response, from being named first to being buried at the bottom of a list.
Agency professionals know that clients will grade themselves against the list, not against the percentage. If they read an answer and see four competitors before their own name, no visibility number is going to fix how that felt. Being ninth in a list of 10 counts as a mention, technically, but almost nobody reads that far. It's the “page 2 of Google” problem in a format with much less room to go around.
Track it per engine, just as you track visibility. Averages do real damage here: a brand that leads on one engine and is missing from another comes out somewhere in the middle, which isn't true of either one.
It's also worth being able to open the full response behind any prompt. Seeing how an answer was built tells you more than the position number ever will, and it's usually where you find the reason for a change.
Here are four of the 30 prompts on our roofing client's list:
Prompt | ChatGPT | AI Overviews | Perplexity |
Best commercial roofing company in Austin | 2nd of 6 | 1st of 5 | 4th of 7 |
Commercial roof replacement cost | 5th of 8 | Not named | Not named |
TPO vs EPDM for a flat roof | Not named | Not named | Not named |
Emergency commercial roof repair Austin | 1st of 4 | 2nd of 6 | Not named |
Read across the rows rather than down the columns for the full story. This client owns emergency repair and local intent, then disappears the moment a question turns technical or gets specific about cost. Average all that into a single position score, and they end up around third, which describes neither the wins nor the gaps.
Give it a quarter before you read much into small movements. Position shifts for reasons that have nothing to do with your client: responses get longer or shorter, models get updated, and the shape of the answer changes.
What to say in the meeting: Set the expectation early that there's no position 1 through 100 here. You're reporting whether your client is named early, named late, or not named at all.
Watch out for presenting this as a rank. The moment a client hears "position," they think keyword rankings, and you've signed up for a comparison that the metric can't support.
The agency that can clearly explain what the numbers actually mean and what they don't earns long-term trust. That level of honesty becomes a true differentiator, especially after clients have been burned by someone who oversold certainty.
Lexi Trimpe, Director of Digital + AI, Franco
Sentiment: how AI engines describe your client
Sentiment in AI search is what the AI says about your client when it names them, including the words it chooses and any warnings it attaches.
This metric has no real equivalent in traditional SEO reporting. A client can be highly visible and still described with information that's two years stale: a product they discontinued, a price that's gone up, a service area they left in 2023.
Being visible with the wrong description does more damage than not being visible at all, and your client would much rather hear about it from you than from a prospect repeating it back to them.
So read the responses. Use the score to sort them, and use the text itself as proof. Keep the full response somewhere you can refer to it a month later.
Report the pattern with the receipt attached:
Recurring description | Appears in | Sentiment | Accurate? | Action |
"Known for fast emergency response" | 9 responses | Positive | Yes | Build on it, add case studies |
"One of the pricier options in the Austin market" | 6 responses | Neutral | Partly | Publish pricing ranges and scope details |
"Primarily serves residential customers" | 4 responses | Neutral | No | Fix site and profile positioning |
Using our roofing company example, the third row is likely the one an agency would lead with. This client stopped taking residential work years ago, so every engine repeating that claim is steering commercial buyers somewhere else.
What to say in the meeting: Lead with the quote, then the pattern. "Three engines describe you as primarily residential, and it's costing us the commercial questions."
Watch out for reporting a sentiment number with no example beside it. Your client can't act on it, and it reads as arbitrary.
There's an upside to this one. When an engine describes a client with stale information, the source is usually something you can reach: their own site, an old profile, a directory listing nobody has touched in years. Wrong information is annoying to find and satisfying to fix.
Citations: which pages an AI answer was built from
Citations in AI search are the URLs an engine pulled from to build its answer, which makes them the clearest map of what to write, refresh, or earn next. Citation share, sometimes called citation rate, is client citations divided by total citations, multiplied by 100.
Start by grouping citations by domain, then sorting them into three buckets; add citation analysis to see which pages AI uses as evidence when mentioning brands:
Pages your client owns
Pages that mention your client
Pages that belong to a competitor
Source quality analysis also helps identify which domains drive the AI engine's confidence in the client. The sorting is what makes the list useful. It's worth tracking which prompts each source feeds, too, since that's the part that turns a list of URLs into an assignment. Over time, citation data also helps you spot shifts in AI citations across prompts and competitors.
Source | Type | Prompts it feeds | Next action |
austinroofingdirectory | Earned | 11 | Update listing, add TPO and EPDM services |
competitorroofing.com | Competitor | 8 | Build a stronger comparison guide |
clientsite.com | Owned | 5 | Add cost ranges and an FAQ block |
reddit.com/r/ | Earned | 4 | Monitor, no action this quarter |
Third-party listicles, directories, and review sites influence citations more than your client's own site. This tells you where the work is actually needed: getting into the roundup, updating the profile, and earning the mention.
The second row is the one that stands out. That competitor-owned page has become the go-to source for the whole category, which is undoubtedly frustrating. It’s also very helpful. Beating the competitor’s ranking for that content is something your team can focus on this quarter.
What to say in the meeting: End every citation slide with actions. "These four sources are answering the questions we want to own, and we’ll focus on shifting two of them this quarter."
Watch out for a list of URLs with nothing attached.
How to track AI search visibility in AgencyAnalytics
Everything above depends on consistency: the same prompts, the same engines, the same location, on the same schedule.
That's no easy task to sustain manually. The prompts live in one spreadsheet, the results in another, and a dozen browser tabs in between.
If a prompt gets reworded, a set gets skipped during a busy week, or someone runs it from a different city, you’ll end up measuring your process rather than your client’s performance.
AgencyAnalytics' AI Tracker handles it for you. The advantage of running this inside AgencyAnalytics is that AI visibility data sits alongside traffic, conversions, and attribution metrics, and it is kept separate and specific to each of your clients.
Here's how it works.
1. Connect AI Tracker as a data source
AI Tracker connects to a client the same way any other integration does, so there's no separate tool to log into and no second place to remember to check.
To get started, open the client you want to track AI content visibility for. On the left sidebar of the dashboard, click Connect Data Source.
Select AI Tracker and then click Enable when prompted.
2. Build the prompt set and group prompts into topics
On the AI Tracker page, click Add Prompts.
Write prompts manually or generate suggestions with AgencyAnalytics’ built-in AI tool, then edit them. Your set is saved to your client instead of being retyped each month.
Tag them the way you'd group keywords, so you can report performance by topic.
Choose the engines you want covered and the location the questions run from. It's a one-time step. AI Tracker monitors visibility across ChatGPT, Google AI Overviews, Google AI Mode, Claude, Perplexity, and Gemini.
3. Watch your credits as you build
Usage updates in real time as you add or remove prompts, and credits are pooled across your whole client roster rather than sold per seat.
Add AI visibility to your client reports
The advantage of running this in AgencyAnalytics is that AI visibility doesn't sit in isolation. When you showcase it in the same report as your client's traffic, conversions, and campaign performance, you can prove ROI.
Add the section to a client report once, and AgencyAnalytics fills it in automatically every month after that.
We all know the AI search visibility question isn't going away. It'll come up again next quarter, and on the next account, and from the client who hasn't thought to ask yet. The idea is to have the answer ready before they do, so the meeting can be about what you're going to do next.
➡️ Try AI Tracker free on AgencyAnalytics for 14 days.
Impress clients and save hours with custom, automated reporting.
Join 7,000+ agencies that create reports in under 30 minutes per client using AgencyAnalytics. Get started for free. No credit card required.
Already have an account?
Log inFrequently asked questions about reporting AI search visibility to clients
AI search visibility is how often and how prominently a brand appears in answers generated by AI engines like ChatGPT, Google AI Overviews, Perplexity, and Gemini, reflecting brand visibility across AI platforms and AI assistants rather than simply whether the brand appears once in a single answer. Where traditional SEO measures a page's position in a list of links, AI search visibility measures whether the brand is named in the answer itself, where it appears in that answer, how it's described, and which sources the engine used to build it.
Most agencies start with 25 to 50 per client. Focus on the questions buyers ask before they know who to call. The exact number matters less than keeping it the same each month, because a percentage only means something when the total it's based on doesn't change. A client with one main service can stay near 25. A client selling several things needs enough prompts to cover each topic without any one of them dominating the set.
You'll see movement in a few weeks and a real trend in a few months. Citations change fastest. Update a directory profile or publish a new comparison page, and it can start showing up in answers within weeks. Position and sentiment take longer, which is why you should give position a full quarter before reading much into small shifts.
Some of it, but not all. If someone clicks a link in an AI answer, that visit appears in your referral traffic under the engine's name. If the answer covers everything and nobody clicks, there's no visit to record. For Google's own AI features, the Generative AI performance report in Search Console shows how your content is performing in AI Overviews and AI Mode. Everything else has to be measured in the answer itself, which is why you track visibility, position, sentiment, and citations instead of guessing from traffic.
Yes, as long as you explain why. Reporting it now gives you a starting point to measure against later. It also catches problems worth fixing right away, like an engine describing your client with outdated information. Put it next to organic traffic, conversions, and leads so your client can see how it connects to the results they care about.
AEO (answer engine optimization) and GEO (generative engine optimization) are two names for the same work: getting a brand mentioned inside AI-generated answers rather than in a list of links. AI search visibility is how you measure that work. Think: if AEO and GEO are the strategy, visibility, position, sentiment, and citations are the scoreboard.
Francois Marchand brings more than 20 years of experience in marketing, journalism, content production, and artificial intelligence. His goal is to equip agency leaders with innovative strategies and actionable advice to succeed in digital marketing, SaaS, and ecommerce.
Read more posts by Francois Marchand


