AI Crawlers and robots.txt: OAI SearchBot, GPTBot, Google and Bing Explained
Understand AI crawlers in 2026, including OAI SearchBot, GPTBot, Google Extended and Bing, and decide what to allow or block in robots.txt.
AI crawler controls are easy to copy and easy to misunderstand. In 2026, the safest approach is to decide separately whether you want content to appear in AI search, whether you want it used for model training and whether a specific crawler is necessary for normal search discovery.
The important point is that not every AI related user agent does the same job.
The three questions to answer before editing robots.txt
Before you add a single rule, decide what outcome you want.
- Do I want my public pages discoverable in search and AI search answers?
- Do I want to allow or disallow model training where the platform offers a separate control?
- Do I have private, paid or low value sections that should not be crawled at all?
Once those decisions are clear, the crawler rules become much easier.
OAI SearchBot and GPTBot are different
OpenAI documents separate crawler controls for search visibility and model training.
OAI SearchBot
OAI SearchBot is used for ChatGPT search features. According to OpenAI's crawler documentation, sites that opt out of OAI SearchBot will not be shown in ChatGPT search answers, although navigational links can still appear in some contexts.
If AI search visibility matters to you, blocking OAI SearchBot can work against that goal.
GPTBot
GPTBot is used for controls related to training OpenAI's generative AI foundation models.
The key detail is that the settings are independent. You can allow OAI SearchBot for search visibility while disallowing GPTBot if that matches your policy.
OpenAI explains these user agents in its official crawler documentation.
What about ChatGPT User?
ChatGPT User is associated with user initiated actions, such as when a user asks ChatGPT to access a page. It is not the same as broad automated crawling.
This distinction matters when you are writing bot policies because blocking one user agent does not necessarily express a policy for all OpenAI product behaviors.
Always check the platform's current documentation before changing production rules because crawler names and behavior can evolve.
What Google Extended controls
Google Extended is often misunderstood as a switch for appearing in Google Search AI features.
It is not simply an “AI Overview crawler.” Google documents it as a control related to the use of site content for certain generative AI model development and grounding outside normal Google Search crawling.
Google's Search guidance says there are no additional technical requirements to appear in AI Overviews or AI Mode beyond normal Search eligibility. A page must be indexed and eligible to appear in Search with a snippet.
So if your goal is Google Search visibility, the most important thing is still whether Googlebot can crawl and index the page under normal Search rules.
See Google's AI features and your website for the current guidance.
Bing and Microsoft AI experiences
Microsoft's AI experiences are closely connected to Bing's web index and Webmaster tools.
In 2026, Bing Webmaster Tools includes an AI Performance report that shows pages cited across supported Microsoft AI experiences, including Copilot and AI generated summaries in Bing.
For site owners, that reinforces an important principle: normal search discoverability and AI visibility are increasingly connected.
A simple robots.txt policy for a public portfolio or marketing site
If the site is entirely public and you want broad discoverability, the simplest policy is usually to allow normal crawlers and avoid unnecessary blocks.
Conceptually:
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Then add platform specific rules only when you have a deliberate reason.
For example, a site owner could choose to allow OpenAI search crawling while disallowing training:
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /
That is only an example policy. Your decision depends on your goals and the platform's current terms.
Do not use robots.txt to protect secrets
robots.txt is a crawling instruction, not an access control system.
Never rely on it to protect:
- Admin pages.
- Customer data.
- Staging dashboards.
- API keys.
- Private documents.
- Paid content that truly requires authentication.
Use authentication, authorization and server side controls for sensitive resources.
A robots.txt file is public. In some cases, listing a private looking path there can actually advertise its existence.
Should you block AI crawlers to protect your content?
There is no one answer for every site.
A publisher may have different priorities from a SaaS company. A personal portfolio trying to build authority may value discoverability more than a site with licensed premium content.
Consider the tradeoff.
Reasons to allow search related AI crawlers
- You want pages to be eligible for AI search citations.
- You want brand discovery through assistants.
- Your content is already public and marketing reach is valuable.
- You are measuring AI visibility as part of SEO.
Reasons to restrict some AI related crawling
- You have licensed or premium content.
- Your company has a specific model training policy.
- Legal or contractual obligations require restrictions.
- A crawler creates unacceptable server load.
The important part is to distinguish search retrieval from training controls when the platform allows it.
How to audit your current robots.txt
Step 1: open the file directly
Visit:
https://yourdomain.com/robots.txt
Make sure it returns a normal text file, not a redirect loop, HTML error page or authentication prompt.
Step 2: look for broad accidental blocks
The most dangerous rule is something like:
User-agent: *
Disallow: /
That blocks all compliant crawlers from the whole site.
Step 3: compare crawler rules with your business goals
For every named AI bot, write down:
- What platform it belongs to.
- What purpose the vendor says it serves.
- Whether you want that purpose.
- The source URL for the current documentation.
Do not make the policy from a copied list with no explanation.
Step 4: check the sitemap reference
Include your XML sitemap URL so crawlers can discover the canonical set of public pages efficiently.
Step 5: verify important pages after changes
After a robots change, confirm that important pages are still crawlable and indexable in Search Console or the relevant webmaster tool.
For OpenAI, the documentation notes that robots updates can take time to be reflected by their systems.
robots.txt and JavaScript rendering
Do not block CSS or JavaScript files that search engines need to render important page content.
Modern websites often depend on client side assets. If the crawler can reach the HTML but cannot render the meaningful content because essential resources are blocked, understanding can suffer.
For critical SEO content, it is still wise to make the important text available directly in the rendered page rather than hiding everything behind interaction.
robots.txt, sitemaps and canonical URLs work together
Think of these tools as different controls:
robots.txt: whether compliant crawlers may fetch certain paths.
XML sitemap: which canonical pages you want crawlers to discover and revisit.
Canonical tag: which URL version should represent duplicate or very similar pages.
Meta robots: whether a crawled page may be indexed or followed, when the crawler can see the directive.
Mixing these concepts causes common mistakes. For example, if you block a page in robots.txt, a crawler may not be able to fetch the page to see its noindex directive.
How AI crawler choices connect to GEO
Crawler access is only the eligibility layer.
Allowing an AI crawler does not guarantee that your site will be mentioned. You still need useful content, a clear entity and enough authority for the system to select your page.
That is why this topic belongs inside a broader GEO strategy and not as a standalone trick.
Likewise, if you want to know whether your visibility changes after crawler policy updates, use a consistent AI visibility tracking framework.
Frequently asked questions
Should I allow OAI SearchBot?
If you want your public website to be eligible to appear in ChatGPT search answers, OpenAI's documentation says OAI SearchBot should not be blocked.
Can I block GPTBot but allow ChatGPT search?
Yes. OpenAI documents OAI SearchBot and GPTBot as separate controls, so a site can allow search crawling while using a different policy for model training.
Does Google Extended control Google AI Overviews?
Google's Search guidance does not describe Google Extended as the requirement for appearing in AI Overviews or AI Mode. Normal Google Search eligibility and indexing remain the foundation.
Does robots.txt prevent someone from opening a URL?
No. It is not a security feature. Use proper access controls for private content.
The takeaway
Do not treat “AI crawler” as one category.
Identify what each user agent actually does, decide what you want, then write the smallest set of rules that matches that policy. Keep important public content crawlable when discovery matters, use true security controls for private resources and revisit platform documentation before making major changes.