
You want to know how to see ai bots on my website, and you want to know it tonight, not after reading a policy debate. Maybe the server bill jumped, the contact form is filling with junk, or a headline made you wonder whether automated software is poking around your pages. The good news is that the first check takes about ten minutes and costs nothing. Your host's access log or a Cloudflare report names the visitors, and the big AI companies publish the names their software uses. The harder part is the next step, because blocking the wrong visitor can take you out of the AI search results you want to appear in. Below, you will find the check in order, a table of the agents you are likely to see, and what each block costs you in visibility.
Key Takeaways
Your host's access log or Cloudflare AI Crawl Control shows which AI agents requested which pages, how often, and what your server answered.
Search crawlers, training crawlers and user-requested agents have different names, different purposes and different controls.
A user-agent string is only a claim, so check important visitors against the operator's published IP ranges before you block or trust them.
It asks compliant crawlers to stay out; firewall and bot-management rules are what enforce a block.
Blocking a search crawler can remove you from that product's answers, while blocking a training crawler does not.
How do I see which AI agents visit my website?
Start where the evidence lives: your access logs, or the bot report in your Cloudflare account if your site sits behind it. Do this before you change a single rule, and save a copy of the logs first. If a visitor turns out to be spoofed or abusive, those records are what let you tell it apart from an ordinary crawler.
- Open the traffic source: In your hosting control panel, find the raw access log or log report. If you use Cloudflare, open AI Crawl Control, which lists crawler names, operators, request totals, unsuccessful requests, robots.txt violations and allow or block controls.
- Filter by name: Search the user-agent field for `OAI-SearchBot`, `GPTBot`, `ChatGPT-User`, `Googlebot`, `ClaudeBot`, `Claude-User`, `PerplexityBot` and `Perplexity-User`.
- Record what each one did: Note the time, the page requested, the status code, the response size, the IP address, the referrer and how often the requests came.
- Check the IP address: Compare it with the operator's published ranges, as described below, before you treat a claimed name as real.
- Read your robots.txt: Open `yourdomain.com/robots.txt` and see which of these names you have already allowed or disallowed, often without realizing it.
- Decide per agent: Allow, slow, challenge or block each one on its own, using the cost table below.
- Test your forms separately: Submit a test message and read your form log, because a crawler in the access log does not prove anything was submitted.
Read the status codes as you go. A 200 means your server delivered the page, a 403 means it refused, a 404 means the page does not exist and a 429 means the visitor made too many requests. A burst of 403 or 429 answers to a search crawler you want is worth fixing the same day.

Which AI agents will I find in my logs?
Each company documents its own agents, and they fall into three groups. Search crawlers find pages so an AI search product can show and cite them. Training crawlers collect pages that may be used to develop models. User-requested agents fetch a page because a person asked for it. OpenAI's bot documentation, Google's crawler list, Anthropic's crawler help article and Perplexity's crawler documentation describe them.
| Name | Company and purpose | Follows robots.txt? | Cost of blocking in AI search |
|---|---|---|---|
| OAI-SearchBot | OpenAI, supports ChatGPT search | Yes, and it can be allowed or disallowed separately from GPTBot | Can keep you out of ChatGPT search answers, though OpenAI says a link may still appear in some cases |
| GPTBot | OpenAI, content that may be used to train models | Yes | Targets training use, not ChatGPT search visibility |
| ChatGPT-User | OpenAI, fetches a page when a user asks | May not, because the request is user initiated | A robots.txt rule may not stop it, so use a firewall rule if you must block it |
| Googlebot | Google Search | Yes, Google says its automated crawlers respect robots.txt | Affects Google Search, including AI Overviews and other Search features |
| Google-Extended | A robots.txt token, not a separate crawler, for certain Gemini training and grounding uses | Yes, it is a robots.txt control | Does not remove you from Google Search; limits certain Gemini uses |
| ClaudeBot | Anthropic, model development | Yes, and Anthropic supports Crawl-delay | Excludes future material from potential training datasets |
| Claude-User | Anthropic, user-requested retrieval | Yes | Can stop Claude from retrieving your pages for a user's answer |
| PerplexityBot | Perplexity, search indexing | Yes | Can remove you from Perplexity search results |
| Perplexity-User | Perplexity, user-requested fetching | Generally ignores it, because a user asked for the fetch | Blocking can prevent user-requested page retrieval |
The pattern matters more than the names. Blocking a training crawler and blocking a search crawler are different decisions, even when the same company runs both. If you want help being found in these products, the post on how to get cited by ChatGPT covers the other side of the same coin, and whether ChatGPT and Google AI can read your website covers what they can see once they arrive.
How can I tell a real AI crawler from a fake one?
A user-agent string is text the visitor chooses to send, so anyone can type "Googlebot" into a script. That is why the IP address matters. Google explains how to verify Googlebot, and the other companies publish IP ranges or a reverse DNS method for their own agents. Cloudflare's bot reference lists the crawlers it recognizes by name and operator.
Do this check whenever a visitor claims a famous name and behaves oddly, for example by requesting thousands of pages in a minute or hitting your login page. If the address belongs to the operator's published ranges, you are looking at the real crawler and can decide on policy. If it does not, the name is a disguise, and a rule built on the name alone will miss it. Use an IP-based rule or your bot-management settings instead.

What did Wikimedia report about an AI agent on October 5, 2026?
On October 5, 2026, the Wikimedia Foundation published a post about activity on its projects that it believed came from OpenAI-operated AI agents. Wikimedia listed what it saw: unauthorized testing edits, potentially malicious edits to a citation-tool configuration, unsuccessful attempts to use its public Etherpad as a proxy for fetching remote data, millions of automated API requests, millions of crawled pages and hundreds of thousands of Wikidata Query Service requests. It said the traffic may have contributed to a partial Wikidata Query Service outage in May. It also said it found no evidence that its systems were used for agent coordination or that its systems or data were compromised, and that it investigated, detected, and reverted or cleaned up the activity. In its words, "almost all of them were testing edits in 'sandbox' areas of the wiki."
For your own site, the useful part is the list of behaviors. Behaviors such as crawling pages and making API requests show up in an access log. Edits and form-style activity show up only in your application logs and form logs. A crawler that fetches public pages is one thing; software that submits, edits or probes is another, and it calls for the form protection described below. Nothing in a log tells you why a visitor behaved the way it did, so judge each one by what it requested and what your server answered.
What can I control, and what does each choice cost in AI search?
You have four levers, and they work at different strengths. Each one has a visibility price, so use the narrowest lever that solves your actual problem.
Robots.txt
Robots.txt is a plain-text file that asks compliant crawlers to stay out of parts of your site. It is documented in RFC 9309, which also makes clear that it is not an access-control mechanism. This example keeps ChatGPT search access open while declining OpenAI's training use:
``` User-agent: OAI-SearchBot Allow: /
User-agent: GPTBot Disallow: / ```
Cost: a Disallow on a search crawler such as OAI-SearchBot or PerplexityBot can remove you from that product's answers, and a Disallow on Googlebot affects Google Search and its AI features. A Disallow on a training crawler such as GPTBot or ClaudeBot does not do that. Google also notes that a page blocked by robots.txt cannot reliably communicate page-level directives such as `noindex` or `nosnippet`, because Google cannot crawl the page to read them. Its robots meta tag guide explains the difference.
Cloudflare AI Crawl Control and managed robots.txt
If your site uses Cloudflare, you can allow or block named AI crawlers from one dashboard, and Cloudflare states that robots.txt compliance is voluntary while AI Crawl Control is the enforcement layer. Its managed robots.txt option writes crawler preferences into your file for you. The cost is the same as above for each name you block, so read the list before you switch a setting on. The earlier post on Cloudflare's Disallow AI Training setting shows how to decline training while keeping Google Search open.
Rate limits
A rate limit caps how many requests a visitor can make in a given time. It protects your server when a crawler is heavy. The cost is that it can also delay or deny a legitimate crawler, so set it high enough that Googlebot and the search crawlers you want can finish their work, and then look at the 429 answers in your log.
Form protection
Contact forms are the one place where automated visits do direct harm. Server-side validation, a hidden honeypot field, a CAPTCHA or managed challenge, spam scoring and monitoring can stop automated submissions without blocking any content crawler. The cost to your AI search visibility is zero, which makes this the first thing to fix if junk submissions are your actual problem.

Does Google Analytics show AI agents?
Not reliably. Many automated requests never run the JavaScript tag that GA4 depends on, so a crawler can fetch a hundred pages without leaving a trace in your reports. GA4 is useful for a different job: counting the people who arrive from an AI service and what they do next. If you are trying to learn whether AI search sends you customers, read the post on whether AI search is bringing you customers. If you are trying to learn who is crawling you, use logs.

Can you tell which AI agents visit your site?
Pick an answer to begin.
1. Where should you look first to see which AI agents requested your pages?
2. You block GPTBot in robots.txt. What happens to ChatGPT search access?
3. Why check an IP address before blocking a visitor that claims to be Googlebot?
Frequently Asked Questions About how to see ai bots on my website
Can Google Analytics identify every AI crawler?
No. Many automated requests do not run the JavaScript that GA4 needs. Use your access logs or a bot-management report.
What is OAI-SearchBot?
It is OpenAI's crawler for surfacing websites in ChatGPT search results. It can be allowed or disallowed separately from GPTBot.
What is GPTBot?
It is OpenAI's crawler for content that may be used to train its generative AI models. Blocking it does not block ChatGPT search.
Does robots.txt enforce a block?
No. It expresses preferences to compliant crawlers. Firewall or bot-management rules enforce an access block.
Will blocking Googlebot affect AI Overviews?
Google says blocking Googlebot affects Google Search and its Search features, including AI features. Its AI features documentation has the details.
Does Perplexity-User follow robots.txt?
Perplexity says user-requested fetching generally ignores robots.txt, because a person asked for the fetch.
Final Thoughts
To see AI agents on your site, open your access log or Cloudflare report, filter for the documented names, verify the addresses and read your robots.txt. Then decide agent by agent: allow the search crawlers you want to be found by, decline training if you prefer, protect your forms, and slow anything that strains your server. Keep a copy of the logs before you change a rule, and write down what you decided.
A site that stays open to useful search systems, resists abusive automation and keeps a record of its decisions is easy to adjust later. If you are comparing ways to spread your visibility, the post on not relying on one site for AI visibility is a useful next read.
You can run the first check yourself. If your logs are unavailable, the traffic is heavy, you run several domains, your firewall rules conflict or a block could cost you revenue, a focused review is worth having, and a website security audit covers the logs and firewall side. If you would like help setting a policy that keeps you visible, Web Leveling offers generative engine optimization, and you can tell us what your logs show and we will help you read them. We work with small and medium businesses across the country and overseas.
Terms
Words you will see in your logs
Tap a term to see what it means.
User agent. A text label a visitor sends with each request to say what software it is; it can be faked.
Crawler. Software that fetches pages automatically so a search or AI product can index or learn from them.
Robots.txt. A text file at the root of your site that asks compliant crawlers which pages to skip.
WAF. A web application firewall, a set of rules that can allow, challenge or block requests.
Rate limit. A cap on how many requests one visitor can make in a set time.
Honeypot. A hidden form field that people never fill in but automated software often does.




