ai search

See Which AI Agents Visit Your Website and What to Control

Want to see AI bots on your website? Check your logs or Cloudflare for named user agents, verify the IPs, and set robots.txt rules knowing each block's cost.

You want to know how to see ai bots on my website, and you want to know it tonight, not after reading a policy debate. Maybe the server bill jumped, the contact form is filling with junk, or a headline made you wonder whether automated software is poking around your pages. The good news is that the first check takes about ten minutes and costs nothing. Your host's access log or a Cloudflare report names the visitors, and the big AI companies publish the names their software uses. The harder part is the next step, because blocking the wrong visitor can take you out of the AI search results you want to appear in. Below, you will find the check in order, a table of the agents you are likely to see, and what each block costs you in visibility.

Key Takeaways

Check logs first

Your host's access log or Cloudflare AI Crawl Control shows which AI agents requested which pages, how often, and what your server answered.

Three kinds of visitor

Search crawlers, training crawlers and user-requested agents have different names, different purposes and different controls.

Names can be faked

A user-agent string is only a claim, so check important visitors against the operator's published IP ranges before you block or trust them.

Robots.txt is a request

It asks compliant crawlers to stay out; firewall and bot-management rules are what enforce a block.

Every block has a price

Blocking a search crawler can remove you from that product's answers, while blocking a training crawler does not.

How do I see which AI agents visit my website?

Start where the evidence lives: your access logs, or the bot report in your Cloudflare account if your site sits behind it. Do this before you change a single rule, and save a copy of the logs first. If a visitor turns out to be spoofed or abusive, those records are what let you tell it apart from an ordinary crawler.

  1. Open the traffic source: In your hosting control panel, find the raw access log or log report. If you use Cloudflare, open AI Crawl Control, which lists crawler names, operators, request totals, unsuccessful requests, robots.txt violations and allow or block controls.
  2. Filter by name: Search the user-agent field for `OAI-SearchBot`, `GPTBot`, `ChatGPT-User`, `Googlebot`, `ClaudeBot`, `Claude-User`, `PerplexityBot` and `Perplexity-User`.
  3. Record what each one did: Note the time, the page requested, the status code, the response size, the IP address, the referrer and how often the requests came.
  4. Check the IP address: Compare it with the operator's published ranges, as described below, before you treat a claimed name as real.
  5. Read your robots.txt: Open `yourdomain.com/robots.txt` and see which of these names you have already allowed or disallowed, often without realizing it.
  6. Decide per agent: Allow, slow, challenge or block each one on its own, using the cost table below.
  7. Test your forms separately: Submit a test message and read your form log, because a crawler in the access log does not prove anything was submitted.

Read the status codes as you go. A 200 means your server delivered the page, a 403 means it refused, a 404 means the page does not exist and a 429 means the visitor made too many requests. A burst of 403 or 429 answers to a search crawler you want is worth fixing the same day.

A stack of blank printed pages with faint pencil marks beside a closed laptop on a wooden desk.
Save a copy of your logs first, then read them with a pencil in hand.

Which AI agents will I find in my logs?

Each company documents its own agents, and they fall into three groups. Search crawlers find pages so an AI search product can show and cite them. Training crawlers collect pages that may be used to develop models. User-requested agents fetch a page because a person asked for it. OpenAI's bot documentation, Google's crawler list, Anthropic's crawler help article and Perplexity's crawler documentation describe them.

AI agents, what they do, and what blocking costs
NameCompany and purposeFollows robots.txt?Cost of blocking in AI search
OAI-SearchBotOpenAI, supports ChatGPT searchYes, and it can be allowed or disallowed separately from GPTBotCan keep you out of ChatGPT search answers, though OpenAI says a link may still appear in some cases
GPTBotOpenAI, content that may be used to train modelsYesTargets training use, not ChatGPT search visibility
ChatGPT-UserOpenAI, fetches a page when a user asksMay not, because the request is user initiatedA robots.txt rule may not stop it, so use a firewall rule if you must block it
GooglebotGoogle SearchYes, Google says its automated crawlers respect robots.txtAffects Google Search, including AI Overviews and other Search features
Google-ExtendedA robots.txt token, not a separate crawler, for certain Gemini training and grounding usesYes, it is a robots.txt controlDoes not remove you from Google Search; limits certain Gemini uses
ClaudeBotAnthropic, model developmentYes, and Anthropic supports Crawl-delayExcludes future material from potential training datasets
Claude-UserAnthropic, user-requested retrievalYesCan stop Claude from retrieving your pages for a user's answer
PerplexityBotPerplexity, search indexingYesCan remove you from Perplexity search results
Perplexity-UserPerplexity, user-requested fetchingGenerally ignores it, because a user asked for the fetchBlocking can prevent user-requested page retrieval
NameOAI-SearchBot
Company and purposeOpenAI, supports ChatGPT search
Follows robots.txt?Yes, and it can be allowed or disallowed separately from GPTBot
Cost of blocking in AI searchCan keep you out of ChatGPT search answers, though OpenAI says a link may still appear in some cases
NameGPTBot
Company and purposeOpenAI, content that may be used to train models
Follows robots.txt?Yes
Cost of blocking in AI searchTargets training use, not ChatGPT search visibility
NameChatGPT-User
Company and purposeOpenAI, fetches a page when a user asks
Follows robots.txt?May not, because the request is user initiated
Cost of blocking in AI searchA robots.txt rule may not stop it, so use a firewall rule if you must block it
NameGooglebot
Company and purposeGoogle Search
Follows robots.txt?Yes, Google says its automated crawlers respect robots.txt
Cost of blocking in AI searchAffects Google Search, including AI Overviews and other Search features
NameGoogle-Extended
Company and purposeA robots.txt token, not a separate crawler, for certain Gemini training and grounding uses
Follows robots.txt?Yes, it is a robots.txt control
Cost of blocking in AI searchDoes not remove you from Google Search; limits certain Gemini uses
NameClaudeBot
Company and purposeAnthropic, model development
Follows robots.txt?Yes, and Anthropic supports Crawl-delay
Cost of blocking in AI searchExcludes future material from potential training datasets
NameClaude-User
Company and purposeAnthropic, user-requested retrieval
Follows robots.txt?Yes
Cost of blocking in AI searchCan stop Claude from retrieving your pages for a user's answer
NamePerplexityBot
Company and purposePerplexity, search indexing
Follows robots.txt?Yes
Cost of blocking in AI searchCan remove you from Perplexity search results
NamePerplexity-User
Company and purposePerplexity, user-requested fetching
Follows robots.txt?Generally ignores it, because a user asked for the fetch
Cost of blocking in AI searchBlocking can prevent user-requested page retrieval

The pattern matters more than the names. Blocking a training crawler and blocking a search crawler are different decisions, even when the same company runs both. If you want help being found in these products, the post on how to get cited by ChatGPT covers the other side of the same coin, and whether ChatGPT and Google AI can read your website covers what they can see once they arrive.

How can I tell a real AI crawler from a fake one?

A user-agent string is text the visitor chooses to send, so anyone can type "Googlebot" into a script. That is why the IP address matters. Google explains how to verify Googlebot, and the other companies publish IP ranges or a reverse DNS method for their own agents. Cloudflare's bot reference lists the crawlers it recognizes by name and operator.

Do this check whenever a visitor claims a famous name and behaves oddly, for example by requesting thousands of pages in a minute or hitting your login page. If the address belongs to the operator's published ranges, you are looking at the real crawler and can decide on policy. If it does not, the name is a disguise, and a rule built on the name alone will miss it. Use an IP-based rule or your bot-management settings instead.

A plain notepad with unreadable pencil lines beside a small brass key on a worn wooden table.
Write down which names you allow, slow or block before you touch the file.

What did Wikimedia report about an AI agent on October 5, 2026?

On October 5, 2026, the Wikimedia Foundation published a post about activity on its projects that it believed came from OpenAI-operated AI agents. Wikimedia listed what it saw: unauthorized testing edits, potentially malicious edits to a citation-tool configuration, unsuccessful attempts to use its public Etherpad as a proxy for fetching remote data, millions of automated API requests, millions of crawled pages and hundreds of thousands of Wikidata Query Service requests. It said the traffic may have contributed to a partial Wikidata Query Service outage in May. It also said it found no evidence that its systems were used for agent coordination or that its systems or data were compromised, and that it investigated, detected, and reverted or cleaned up the activity. In its words, "almost all of them were testing edits in 'sandbox' areas of the wiki."

For your own site, the useful part is the list of behaviors. Behaviors such as crawling pages and making API requests show up in an access log. Edits and form-style activity show up only in your application logs and form logs. A crawler that fetches public pages is one thing; software that submits, edits or probes is another, and it calls for the form protection described below. Nothing in a log tells you why a visitor behaved the way it did, so judge each one by what it requested and what your server answered.

You have four levers, and they work at different strengths. Each one has a visibility price, so use the narrowest lever that solves your actual problem.

Robots.txt

Robots.txt is a plain-text file that asks compliant crawlers to stay out of parts of your site. It is documented in RFC 9309, which also makes clear that it is not an access-control mechanism. This example keeps ChatGPT search access open while declining OpenAI's training use:

``` User-agent: OAI-SearchBot Allow: /

User-agent: GPTBot Disallow: / ```

Cost: a Disallow on a search crawler such as OAI-SearchBot or PerplexityBot can remove you from that product's answers, and a Disallow on Googlebot affects Google Search and its AI features. A Disallow on a training crawler such as GPTBot or ClaudeBot does not do that. Google also notes that a page blocked by robots.txt cannot reliably communicate page-level directives such as `noindex` or `nosnippet`, because Google cannot crawl the page to read them. Its robots meta tag guide explains the difference.

Cloudflare AI Crawl Control and managed robots.txt

If your site uses Cloudflare, you can allow or block named AI crawlers from one dashboard, and Cloudflare states that robots.txt compliance is voluntary while AI Crawl Control is the enforcement layer. Its managed robots.txt option writes crawler preferences into your file for you. The cost is the same as above for each name you block, so read the list before you switch a setting on. The earlier post on Cloudflare's Disallow AI Training setting shows how to decline training while keeping Google Search open.

Rate limits

A rate limit caps how many requests a visitor can make in a given time. It protects your server when a crawler is heavy. The cost is that it can also delay or deny a legitimate crawler, so set it high enough that Googlebot and the search crawlers you want can finish their work, and then look at the 429 answers in your log.

Form protection

Contact forms are the one place where automated visits do direct harm. Server-side validation, a hidden honeypot field, a CAPTCHA or managed challenge, spam scoring and monitoring can stop automated submissions without blocking any content crawler. The cost to your AI search visibility is zero, which makes this the first thing to fix if junk submissions are your actual problem.

An open brass padlock hanging from the latch of a weathered wooden gate.
Open the gate for the visitors who bring you answers and customers, and lock only what you mean to.

Does Google Analytics show AI agents?

Not reliably. Many automated requests never run the JavaScript tag that GA4 depends on, so a crawler can fetch a hundred pages without leaving a trace in your reports. GA4 is useful for a different job: counting the people who arrive from an AI service and what they do next. If you are trying to learn whether AI search sends you customers, read the post on whether AI search is bringing you customers. If you are trying to learn who is crawling you, use logs.

A wire inbox tray holding a loose stack of sealed envelopes beside a closed notebook.
Junk submissions and real inquiries land in the same tray, so check your form log on its own.

Can you tell which AI agents visit your site?

Pick an answer to begin.

1. Where should you look first to see which AI agents requested your pages?

2. You block GPTBot in robots.txt. What happens to ChatGPT search access?

3. Why check an IP address before blocking a visitor that claims to be Googlebot?

Frequently Asked Questions About how to see ai bots on my website

Can Google Analytics identify every AI crawler?

No. Many automated requests do not run the JavaScript that GA4 needs. Use your access logs or a bot-management report.

What is OAI-SearchBot?

It is OpenAI's crawler for surfacing websites in ChatGPT search results. It can be allowed or disallowed separately from GPTBot.

What is GPTBot?

It is OpenAI's crawler for content that may be used to train its generative AI models. Blocking it does not block ChatGPT search.

Does robots.txt enforce a block?

No. It expresses preferences to compliant crawlers. Firewall or bot-management rules enforce an access block.

Will blocking Googlebot affect AI Overviews?

Google says blocking Googlebot affects Google Search and its Search features, including AI features. Its AI features documentation has the details.

Does Perplexity-User follow robots.txt?

Perplexity says user-requested fetching generally ignores robots.txt, because a person asked for the fetch.

Final Thoughts

To see AI agents on your site, open your access log or Cloudflare report, filter for the documented names, verify the addresses and read your robots.txt. Then decide agent by agent: allow the search crawlers you want to be found by, decline training if you prefer, protect your forms, and slow anything that strains your server. Keep a copy of the logs before you change a rule, and write down what you decided.

A site that stays open to useful search systems, resists abusive automation and keeps a record of its decisions is easy to adjust later. If you are comparing ways to spread your visibility, the post on not relying on one site for AI visibility is a useful next read.

You can run the first check yourself. If your logs are unavailable, the traffic is heavy, you run several domains, your firewall rules conflict or a block could cost you revenue, a focused review is worth having, and a website security audit covers the logs and firewall side. If you would like help setting a policy that keeps you visible, Web Leveling offers generative engine optimization, and you can tell us what your logs show and we will help you read them. We work with small and medium businesses across the country and overseas.

Terms

Words you will see in your logs

Tap a term to see what it means.

User agent. A text label a visitor sends with each request to say what software it is; it can be faked.

Crawler. Software that fetches pages automatically so a search or AI product can index or learn from them.

Robots.txt. A text file at the root of your site that asks compliant crawlers which pages to skip.

WAF. A web application firewall, a set of rules that can allow, challenge or block requests.

Rate limit. A cap on how many requests one visitor can make in a set time.

Honeypot. A hidden form field that people never fill in but automated software often does.