To track AI traffic in GA4, you build a custom channel group that matches the referrer hostnames of AI assistants: chatgpt.com, perplexity.ai, claude.ai, copilot.microsoft.com, gemini.google.com and their relatives, then cross-read it against Search Console and your server logs. That gets you the visible slice. The larger, uncomfortable truth is that most AI-influenced traffic arrives with no referrer at all and lands in Direct, and no configuration fixes that.
Key Takeaways
- AI assistant referrals are visible in GA4 when the assistant passes a referrer. Build one custom channel group with a regex covering all major hostnames rather than reading a scattered Referral report.
- Chat assistants and their search products use different hostnames (
chatgpt.comvsopenai.com,gemini.google.comvsgoogle.com). Catch the family, not the single domain. - Search Console reports impressions and clicks from Google's AI surfaces inside existing Search performance data rather than as a fully separate, cleanly split channel: read it as directional, not as a clean AI number.
- Server logs are the only place you see AI crawlers (GPTBot, PerplexityBot, ClaudeBot, and others). Crawl activity tells you whether you're eligible to be cited; it does not tell you whether you were.
- Everything above measures the minority. Zero-click answers, cross-device journeys, and referrer-stripped app traffic mean the majority of AI influence lands in Direct. Self-reported attribution is the only instrument that reaches it.
Why the Existing Advice Is Half a Stack
Search for how to track AI traffic and you get two disconnected genres. One is "add this regex to GA4," written by analytics people who never mention Search Console. The other is "Google now reports AI surface data," written by SEOs who never mention GA4 channel groups or log files. Neither tells you what the combined picture looks like, and neither tells you honestly how much of the picture is missing.
This post is the whole stack in one place, ending with what it still can't see.
Layer 1: The Referrer Hostname List
When someone clicks a citation inside an AI assistant, your server usually receives an HTTP referrer identifying the assistant's domain. That's the raw signal underneath everything else.
Here's the working list, with what each source actually tells you.
| AI platform | Referrer hostname(s) | What you can see | What you can't |
|---|---|---|---|
| ChatGPT | chatgpt.com, openai.com | Sessions, landing page, on-site behaviour, key events | The prompt, whether you were cited but not clicked, which conversation turn |
| Perplexity | perplexity.ai, www.perplexity.ai | Sessions and landing page; often the highest click-through of the group | Query text, citation position |
| Claude | claude.ai, anthropic.com | Sessions and landing page | Prompt, citation-without-click |
| Microsoft Copilot | copilot.microsoft.com, bing.com (Copilot-embedded results) | Sessions; Bing-blended traffic is hard to separate cleanly | Whether the click came from a Copilot answer or a classic Bing result |
| Google Gemini | gemini.google.com, bard.google.com (legacy) | Sessions from the standalone Gemini app | Anything from AI Overviews inside google.com, that's indistinguishable from organic search in the referrer |
| Google AI Mode / AI Overviews | google.com | Nothing distinct in GA4, it arrives as organic search | The AI-surface split; Search Console is the only partial view |
| Meta AI | meta.ai | Sessions | Prompt, in-app journeys |
| Grok / x.ai | grok.com, x.ai | Sessions | Prompt |
| DeepSeek | chat.deepseek.com, deepseek.com | Sessions | Prompt |
| Mistral Le Chat | chat.mistral.ai | Sessions | Prompt |
| You.com | you.com | Sessions | Prompt |
| Poe | poe.com | Sessions | Which underlying model answered |
Two things to internalise from that table. First, the "what you can't see" column is longer than the "can" column in every row. Second, the single largest AI surface by volume for most sites, Google's AI Overviews, produces no distinguishing referrer at all. It looks like organic search, because it is organic search, delivered differently.
The hostname list will go stale
New assistants launch, domains change, products get folded into others. Whatever regex you build, revisit it quarterly. Sort your Referral report by session count and scan for hostnames you don't recognise, that's how you find the next one before anyone writes a blog post about it.
Layer 2: The GA4 Custom Channel Group
Custom channel groups live in GA4's Admin area under the data display settings. You create a new group, add a channel above the defaults, and give it a condition set. Conditions accept regex matching on the source, which is what makes one rule cover a dozen domains. Google's interface for this has been reorganised more than once, check Google Analytics Help for the current path rather than following a screenshot.
The copy-paste regex
Set your condition to Source matches regex:
^(chatgpt\.com|.*\.chatgpt\.com|openai\.com|.*\.openai\.com|perplexity\.ai|.*\.perplexity\.ai|claude\.ai|.*\.claude\.ai|anthropic\.com|copilot\.microsoft\.com|gemini\.google\.com|bard\.google\.com|meta\.ai|grok\.com|x\.ai|chat\.deepseek\.com|deepseek\.com|chat\.mistral\.ai|you\.com|poe\.com|phind\.com|iask\.ai|andisearch\.com)$Name the channel AI Assistants. Order it above Referral and Organic Search, because channel groups evaluate top-down and the first matching rule wins, put it below Referral and every one of these will be swallowed as a generic referral.
Two implementation notes that trip people up
Custom channel groups apply to data going forward plus a bounded window of historical data. They are not a full retroactive rewrite of your history. Create the group before you need the report.
And note the deliberate omission: bing.com is not in that regex, even though Copilot-embedded results can arrive from it. Including it would drag your entire Bing organic footprint into the AI bucket and make the channel meaningless. If Copilot matters to you specifically, build it as a separate channel and accept the ambiguity explicitly rather than hiding it.
Building it as a segment instead
If you don't want to touch channel groups yet, the same regex works as an exploration segment. In an exploration, create a session-scoped segment where Session source matches the regex above. You get the same population without changing a property-level setting, useful for testing before you commit.
The segment approach is also better for behavioural analysis. Once you have it, compare AI-referred sessions against organic search sessions on engagement rate, pages per session, and key event rate. In most content sites I've looked at, AI-referred visitors arrive further down the funnel and convert at a higher rate on a much smaller base. That ratio is the real story, and it's the number worth reporting.
Layer 3: Search Console and Google's AI Surfaces
This is where honest reporting matters most, because it's where the most confident nonsense gets published.
Google has stated that clicks and impressions from AI experiences, AI Overviews and AI Mode, are included in Search Console's Performance report data. What Google has not given you is a clean, complete, permanently-labelled split showing "here is exactly your AI Overview performance, isolated." The reporting granularity here has changed over time and may change again; Google Search Central is the authoritative source and its documentation is where you should verify current behaviour before you build a report on it.
How to read it honestly
What you can do reliably:
- Track impression-to-click ratio over time on your informational queries. If impressions hold steady or grow while clicks fall, that pattern is consistent with answers being satisfied on the results page. It is evidence, not proof.
- Segment by query type. Definitional and "what is / how does" queries are most exposed to AI answer surfaces. Comparison, pricing, and branded queries are far less so. Watch those cohorts separately, a blended sitewide CTR number will hide the entire effect.
- Watch branded query volume. If AI assistants are naming your brand, the downstream effect is people searching your brand name. Rising branded impressions alongside flat referral traffic is one of the few positive AI signals you can actually observe.
What you cannot do: produce a number labelled "AI Overviews sent us X sessions." That number does not exist in your tooling. If a vendor dashboard shows it to you, ask precisely how it was derived: it's a model, not a measurement. Search Engine Journal and Search Engine Land have both tracked the evolving reporting picture closely and are worth reading for changes.
Layer 4: AI Crawlers in Server Logs
Referrers tell you about people. Logs tell you about machines, and the machines decide whether you're eligible to be cited at all.
The crawler user-agents worth grepping for:
GPTBot, OpenAI's training crawlerOAI-SearchBot, OpenAI's search index crawlerChatGPT-User, live retrieval when a user's prompt triggers a fetchPerplexityBot, Perplexity's index crawlerClaudeBotandClaude-User, Anthropic's crawlersGoogle-Extended, Google's AI training control tokenBingbot, still the crawler behind Microsoft's AI surfacesApplebot-Extended, Apple's AI training controlBytespider,Amazonbot,CCBot, others in the mix
What to actually do with this
Filter your access logs by user-agent and answer four questions: which of these are crawling you at all, how often, which URLs they hit, and, critically, what status codes they receive. A section of your site returning 403 to PerplexityBot while returning 200 to Googlebot is a self-inflicted wound, and it's more common than you'd expect because bot-blocking rules in CDN WAF configs are often set once and never reviewed.
Also verify your robots.txt says what you intend. Some teams block AI crawlers deliberately and correctly. Many block them accidentally by copying a rule set from somewhere else, then wonder why they're never cited.
Live retrieval agents like ChatGPT-User are the most interesting signal, because they fire when a real user's prompt triggered a fetch of your page. Frequency there is the closest thing to a real-time citation indicator you'll get from logs.
What logs cannot tell you
Being crawled is not being cited. Being cited is not being clicked. Logs give you the top of that chain only. Treat crawl frequency as an eligibility metric, not a performance metric.
Layer 5: What You Still Cannot Measure
Here's the section every other post on this topic leaves out.
Zero-click answers
The dominant AI use case is the one where nobody visits your site. Your content is read, synthesised, and delivered inside the assistant. You get brand exposure and no session. There is no tracking configuration that captures this, because there is no HTTP request to your server to capture. The only proxies are branded search volume and self-reported attribution.
The no-referrer majority
A large share of AI-influenced visits arrive with nothing attached. Reasons stack up:
- Assistants in native mobile apps and desktop clients often don't pass a web referrer
- Users read an answer on a phone, then search or type your URL on a laptop later
- Users copy a URL out of an answer and paste it directly
- Privacy settings and referrer policies strip the header
- The visit happens hours or days after the AI conversation
All of these land in Direct. Your carefully-built AI Assistants channel captures the referrer-preserving minority.
The delayed-influence problem
Someone asks an assistant which coding bootcamps in India have strong placement records. Your brand is named. They do nothing that day. Eleven days later they search your brand and apply. GA4 records organic search, branded. That is a correct record of the final click and a completely misleading record of what caused the application.
Why this matters more for organic-led brands
If you buy traffic, you have click IDs and platform-side reporting. If you earn traffic, you have referrers, and AI surfaces are systematically eroding the referrer. The measurement gap is not evenly distributed; it falls hardest on exactly the teams doing the work that AI systems reward.
When I was growing Masai School's social audience, Instagram from 26K to 117K, LinkedIn from 50K to 160K, the same structural problem applied to social. The channel doing the work was the channel least able to prove it. AI search is that problem again, one layer deeper.
The Instrument That Closes the Gap
One question on your lead form: "How did you hear about us?"
When people write "ChatGPT recommended you," "found you through Perplexity," or "an AI tool mentioned your name," you have a data point your entire tracking stack could not produce. It's self-selected, incomplete, and biased: and it's still the single highest-signal input you have about AI-driven discovery. I've made the full argument, including how to implement and reconcile it, in Self-Reported Attribution Beats Your Tracking Stack.
Putting the Stack Together
A realistic monthly review takes twenty minutes:
- GA4 AI Assistants channel: sessions, key events, engagement rate vs organic search. Direction, not precision.
- Search Console: impressions vs clicks on informational query cohorts, and branded query trend.
- Server logs: which AI crawlers hit you, how often, what status codes came back.
- Self-reported attribution, the count and text of AI mentions in your "how did you hear about us" responses.
- Direct traffic trend, rising Direct alongside rising AI crawl activity and rising branded search is the composite signature of AI influence you can't see directly.
Report all five. Report the fifth as inference, clearly labelled. Nobody is fooled by false precision, and your credibility is worth more than a confident-looking number.
FAQ
How do I see AI traffic in GA4 right now, without configuring anything?
Open Reports, go to Traffic acquisition, switch the dimension to Session source, and search for chatgpt, perplexity, claude, copilot, or gemini. You'll see whatever is arriving with a referrer. That's your five-minute answer; the channel group makes it repeatable.
Does ChatGPT pass a referrer?
Web ChatGPT generally does: sessions from citation clicks typically show chatgpt.com as the source. Native mobile and desktop app traffic is far less reliable, and traffic from copied-and-pasted links carries nothing at all.
Can I track AI Overviews traffic separately in GA4?
No. AI Overviews appear inside Google Search results, so clicks arrive with google.com as the referrer and are indistinguishable from any other organic click. Search Console's Performance data includes AI surface activity but doesn't hand you a clean isolated split. Anyone claiming a precise AI Overviews session count in GA4 is modelling, not measuring.
Is AI traffic worth tracking if the volume is tiny?
Yes, for two reasons. The trend line matters more than the absolute number: a channel going from 0.3% to 2% of sessions in six months is a directional signal worth acting on. And AI-referred visitors typically arrive with unusually high intent, so their key event rate is often well above your site average. Small volume, disproportionate value.
Should I block AI crawlers in robots.txt?
Depends on whether you want citations. Blocking training crawlers while allowing search and live-retrieval crawlers is a defensible middle position. Blocking everything means you won't be cited. Make it a deliberate decision and document it, rather than inheriting whatever your CDN's default bot rules do.
Why is my Direct traffic growing so fast?
Some of it is genuinely AI influence: assistants recommend you, people arrive without a referrer, GA4 files it as Direct. Some of it is untagged organic social, in-app browsers stripping referrers, and dark social sharing. All of these are worth separating, which is exactly what a "how did you hear about us" field does and analytics cannot.
Do I need a paid tool to track AI visibility?
No, for the measurement described here: GA4, Search Console, and log access cover all of it. Paid tools add prompt-level visibility tracking (are you named in answers to specific questions), which is genuinely useful and genuinely a different thing from traffic measurement. Get the free stack right before you buy the paid one.
What's the difference between AI traffic and AEO or GEO?
AI traffic is measurement: how many sessions arrived from AI surfaces. AEO and GEO are practice: how you structure content so AI systems can extract, cite, and recommend it. This post is about the first. Doing the second is what generates the traffic you then try to measure.
How often should I update my AI referrer regex?
Quarterly. Scan your Referral report for unfamiliar hostnames with meaningful session counts. New assistants appear faster than blog posts get updated, including this one.
Can I use UTMs to track AI traffic?
Only for links you control and place yourself. You can't add UTM parameters to a citation an AI system generates from your organic content. Referrer-based tracking is the only mechanism available for AI referrals.
I write about organic growth measurement for Indian edtech and startup brands, including the AEO and GEO work that produces the AI visibility this post tries to measure. More at younusfardeen.com if you want to see how I'd build this for your stack.