How to Check Whether Your Website Can Appear in ChatGPT Search: A 6-Step Publisher Audit

Getting a page indexed by a traditional search engine does not automatically mean it is eligible to appear as a cited source in ChatGPT Search. OpenAI uses a dedicated search crawler, OAI-SearchBot, and gives publishers separate controls for search visibility and model-training preferences.

This six-step audit helps publishers check the parts they can actually control: crawl access, indexing signals, bot separation, infrastructure blocks, and referral measurement.

1. Confirm the page is public and crawlable

Start with the basics. The page should be publicly reachable without a login, and your site should not unintentionally block the paths you want discovered. OpenAI says public websites can appear in ChatGPT search, but inclusion in summaries and snippets depends in part on whether OAI-SearchBot can access the content.

2. Check robots.txt for OAI-SearchBot

OpenAI identifies OAI-SearchBot as the crawler used to surface websites in ChatGPT search features. If you want eligible pages to be discoverable, make sure your robots.txt rules do not block that user agent.

User-agent: OAI-SearchBot
Allow: /

OpenAI notes that changes to robots.txt can take roughly 24 hours to be reflected in its search systems, so do not expect an immediate result after editing the file.

3. Keep search visibility separate from training controls

OAI-SearchBot and GPTBot serve different purposes, and OpenAI documents their controls as independent. A publisher can allow OAI-SearchBot for search discovery while disallowing GPTBot for potential model-training use. Do not assume that changing one setting automatically changes the other.

4. Review noindex and page-level signals

If you do not want a page surfaced, OpenAI’s publisher guidance points to the noindex meta tag. There is an important sequencing detail: a crawler has to be allowed to access the page in order to read that meta tag.

That means a blanket robots.txt block and a page-level noindex directive are not interchangeable controls. Use them deliberately rather than assuming they communicate the same instruction.

5. Check CDN, firewall, and bot-protection rules

A correct robots.txt file is not enough if a CDN, firewall, or automated bot-defense layer returns a 403 response to the crawler. OpenAI recommends checking these infrastructure controls and, where appropriate, allowing traffic associated with its published crawler user agents and IP ranges.

This is especially relevant for sites using aggressive anti-bot protection, because a valid crawler can still be blocked before it ever reaches the page.

6. Measure ChatGPT referral traffic instead of guessing

Visibility should be measured with observable data rather than inferred from a few manual searches. OpenAI says referral URLs from ChatGPT include utm_source=chatgpt.com, which publishers can track in analytics tools such as Google Analytics.

Use that referral signal as evidence of visits from ChatGPT search, but do not confuse referral traffic with citation quality, ranking, or total AI visibility. A page can be technically crawlable and still receive little or no traffic.

A compact ChatGPT Search visibility checklist

  1. Public: Is the target page reachable without authentication?
  2. Robots: Is OAI-SearchBot allowed on the relevant path?
  3. Controls: Are OAI-SearchBot and GPTBot configured intentionally and separately?
  4. Indexing: Are noindex directives consistent with what you want surfaced?
  5. Infrastructure: Is a CDN, WAF, or bot-defense rule blocking the crawler?
  6. Measurement: Are you tracking referrals tagged with utm_source=chatgpt.com?

The key rule: treat AI-search visibility as a crawlability-and-measurement problem first, not as a promise that any particular page will be cited.

Sources

Fact-check date: September 27, 2026.

Choose language: English · 한국어 · 日本語 · Deutsch