OAI-SearchBot vs GPTBot: WordPress Robots.txt Setup

A WordPress crawler policy is useful when its saved settings match its public response. This OAI-SearchBot vs GPTBot tutorial guides you through the policy choice, deployment, and verification, including checks that a dashboard setting alone cannot provide.

This tutorial shows how to choose a policy, apply it through the correct WordPress robots.txt source, and verify what visitors and crawlers actually receive. The goal is a working access configuration with evidence you can review.

Configuration guidance checked against official documentation on 10 October 2026.

Quick Answer

Use separate robots.txt groups for OAI-SearchBot and GPTBot to permit search crawling and disallow training crawling. Check the delivered file, matching path rules, and firewall responses. This verifies access; citations require separate observation.

Key Takeaways

  • Choose a policy before changing the file, and record who maintains it.
  • Edit the source that serves the public robots.txt response.
  • Review existing groups before adding another rule for the same bot.
  • Carry required path restrictions into a bot-specific group.
  • Verify the response after clearing relevant caches.
  • Keep crawl evidence, answer citations, and visitor analytics as separate measurements.

Table of Contents

OAI-SearchBot vs GPTBot: Which Setting Controls What?

AgentPurposeConfiguration implication
OAI-SearchBotSearch discovery for ChatGPTUse its group for search-crawl policy
GPTBotCrawling content that may support model trainingSet training-crawl policy independently
ChatGPT-UserSome user-requested page visitsRobots.txt may not govern these requests; it is not the search inclusion control

The official OpenAI crawler documentation describes these independent controls. Search opt-out excludes search answers, though navigational links can remain.

Keep the implementation focused on your intended public content. For the broader purpose of crawl instructions, read our robots.txt beginner guide. If you want to understand citation-oriented content work after access is checked, our GEO guide explains the wider context.

What Do You Need Before Making Changes?

Prepare the following before editing production settings:

  • The correct website host: Know which public hostname serves your articles and whether other hosts redirect to it.
  • The current delivered file: Open the site’s root-level /robots.txt, copy its complete contents, and record when you retrieved it.
  • Editing access: Have WordPress administrator access to the relevant plugin settings, or access to the host’s file manager for a physical file.
  • Cache access: Identify any page cache, server cache, or content delivery network that can retain the old response.
  • Security visibility: Know where to inspect firewall or bot-management events, or which hosting administrator can provide them.
  • A small test set: Choose a public article, an important page, and one path that should stay excluded.
  • A recovery copy: Save the previous complete file and the current security-rule settings so a mistaken change can be reversed.

Decide the policy in plain language. For example: “Public learning articles should be reachable for search; training crawling should be disallowed.” If the site owner wants a different policy, write that down first. Access configuration is a publishing decision, not a requirement to accept every automated request.

Check for existing entries for both bot names, a wildcard group, and any blanket exclusions. Include rules managed by plugins or the hosting platform in your review. Moving a restriction into the wrong group can produce a different policy from the one you intended.

A Robots.txt Configuration for Search Access and Training Opt-Out

The following is an illustrative configuration for a root-level WordPress installation. It preserves a familiar administrative-path exclusion, adds an explicit search-crawler group, and applies the chosen training opt-out. Adapt it to your actual paths and existing rules before use.

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

User-agent: OAI-SearchBot
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

User-agent: GPTBot
Disallow: /

This sample is not a replacement for every site’s complete file. Retain your verified sitemap declaration and any other intentional restrictions. If WordPress is installed under a subdirectory, review its actual administrative paths instead of assuming the sample matches.

Under the Robots Exclusion Protocol, matching bot groups take precedence over the wildcard fallback, and repeated matching groups are combined. The most specific matching path rule applies. Unmatched paths are allowed. Consequently, required exclusions should be included in the specific bot group instead of relying on inheritance from User-agent: *.

Consider a public article path and the administrative path in the sample. The article has no matching exclusion in the search group, while the administrative directory does. The longer exception permits the listed AJAX endpoint. The training group excludes every path because its rule begins at the site root.

When reviewing a real file, examine all matching groups. Adding an allow rule at the bottom is not a general way to cancel an earlier restriction. A more specific exclusion can still apply, and two competing entries can make future maintenance difficult.

Keep confidential information behind actual authentication. The protocol describes crawl instructions, not access authorization; listing a sensitive path in a publicly readable file can expose its name.

How to Apply the Policy in WordPress

Step 1: Identify the Source of the Public Response

WordPress can generate robots.txt output dynamically. Its do_robots function produces the default response, and the robots_txt filter lets code modify that output.

A physical file or a response generated by the hosting or security layer may be what the public URL actually serves. Inspect the site’s document root through the host’s file manager when necessary. Do not assume the file must exist merely because the URL works.

Select one maintained source of truth. If a physical file controls delivery, edit that file. If the response is virtual and managed by an SEO plugin, use that plugin’s editor. If you cannot identify the source, ask the host to confirm how the robots.txt request is routed before making changes.

Step 2: Use the Appropriate Editor

For a Rank Math-managed virtual file, the current Rank Math instructions use this route:

  1. Open WordPress Dashboard → Rank Math SEO.
  2. Enable Advanced Mode, then open General Settings → Edit robots.txt.
  3. Review the existing contents and reconcile any matching bot groups.
  4. Apply the intended policy and select Save Changes.

The free editor supports this task. Rank Math’s documentation explains that an existing physical file prevents its virtual editor from controlling the served file. If a physical file exists, keep using it or plan a deliberate migration; preserve its contents and verify the replacement before removing the old source.

For a host-managed physical file, locate the website’s actual document root and edit the file named robots.txt as plain text. Keep the verified existing rules, add or revise the applicable groups, and save. Do not create a WordPress post titled “robots.txt”; an ordinary content page is not the root response crawlers request.

Developer-maintained sites can use WordPress’s documented filter instead. That route needs the same review of existing groups and delivered output. Adding code is unnecessary when an existing editor already manages the correct response.

Step 3: Clear Relevant Caches and Fetch Again

After saving, clear the cache layers that can serve this response. Depending on the installation, that may include the WordPress cache, a server cache, and a CDN cache for the robots.txt URL.

Reopen the public file in a logged-out browser session and compare it with the intended text. Check its HTTP response as well: a cache can retain old text even after a successful save. If the delivered response differs, resolve the source or cache issue before changing the policy again.

How to Verify the Delivered Policy and Page Access

Check the HTTP Response and Text

If you are comfortable with a terminal, this is a read-only check using a reserved example domain:

curl -i https://example.com/robots.txt

Replace the example domain with your actual public hostname. Inspect the final response, status, content type, and text. A normal successful response should contain the intended plain-text rules, rather than a login form, challenge page, or branded error document.

The root location, plain-text format, and host scope are described in Google’s robots.txt specification documentation. Check important alternate hosts separately, including their redirects. Do not assume a rule retrieved from one hostname proves what another hostname serves.

Record the check time and the response you saw. Avoid depending only on an editor screenshot or a local copy of the file. Those cannot show whether an old version remains at the edge of the network.

Test Representative Paths Against the Matching Group

Read all matching rules for each test path. Use a robots parser or tester that lets you specify the bot token, then compare its result with a manual review. Record the tool, token, path, and result. A tester designed only for Googlebot cannot establish how another crawler actually behaves.

Test itemExpected result for the illustrative policyWhat the check establishes
Public article with the search bot tokenNo matching exclusionThe policy permits that path
Administrative directory with the search bot tokenMatching exclusionThe intended directory restriction remains
Public article with the training bot tokenMatching root exclusionThe training opt-out rule is present and applies
Live public article requestUseful page content is returnedThe server can deliver the page to that requester

Policy permission and successful retrieval are separate checks. A parser can say “allowed” while the firewall still rejects the request. Conversely, a page can load for your browser even though the policy tells an automated crawler not to fetch it.

Investigate Firewall and Bot-Management Responses

Review relevant security events for the public article and robots.txt path. Look for blocks, rate limits, or browser challenges. Record which rule fired and which layer produced the response.

A request claiming a bot name is not sufficient proof of identity. Use the current crawler-specific IP information linked from OpenAI’s crawler reference when evaluating genuine requests.

Ask the hosting or security administrator to review the specific rule responsible for a confirmed block. Any exception should be limited to the intended crawler and public resources. Preserve protections for login, administration, and unrelated traffic, then verify a subsequent real request.

A manually changed user-agent is useful for finding responses that vary by header, but it does not turn your request into the real crawler. Your IP, request behavior, and network location can produce a different result. Label such a test as a diagnostic simulation.

WordPress crawler setup process: choose policy, edit the source, refresh caches, test paths, and check requests.
Verify the delivered policy and actual retrieval as separate checks.

An Illustrative WordPress Repair Example

Imagine a fictional educational website with a virtual robots.txt editor, a CDN, and a bot-management rule. The owner wants public tutorials reachable under the sample policy. The following observations and repairs are illustrative; they are not results from testing a real website.

ObservationLikely layer to inspectPriority and actionEvidence needed to close the issue
The editor contains the new policy, but the public file shows old textCDN or server cache; response sourceFirst: confirm the active source and refresh the cached fileA fresh external response matches the saved policy
A second matching bot group contains an unintended exclusionCombined robots.txt rulesFirst: reconcile the duplicate groups and preserve intended restrictionsPath checks agree with the chosen policy
A verified crawler request receives a challenge on a public tutorialSecurity rule or bot-management policyNext: review the specific matching rule with the administratorA subsequent verified request retrieves the intended page
There are no recorded citations after delivery is repairedAnswer sampling and content relevanceLater: review useful content and sample relevant questionsRecorded observations of actual answers, without predicting inclusion

Start with delivery because an undelivered policy makes further rule edits difficult to interpret. Then confirm the rules that apply to the chosen paths. Investigate retrieval blocks after those checks, and evaluate content visibility as a separate task.

The successful outcome of this repair is a documented, consistent access configuration. Whether a search answer later uses a page requires separate observation. Keep that distinction in the repair register instead of closing the ticket with an assumed traffic increase.

Troubleshooting: Why Does the Configuration Still Look Wrong?

The Saved File Does Not Match the Public File

Check for a physical file, a plugin-managed virtual response, a server override, and stale cache content. Compare the complete response at each accessible layer. Ask the host to identify the active source if necessary; repeatedly saving the same plugin setting will not resolve a different source serving the URL.

A Tester Says a Public Path Is Disallowed

Inspect every matching bot group and the exact path being tested. Look for an unintended root exclusion or a longer matching path rule. Confirm that the tester is using the correct token and the latest delivered file. Correct the specific conflict, then rerun the same test set.

The File Is Allowed, but the Page Returns an Error

Separate policy parsing from HTTP retrieval. Investigate the public page’s status, redirects, authentication, security challenge, and rate limit. A working robots.txt endpoint cannot repair a broken article URL. Our guide to 404 errors covers missing-page diagnosis.

A Change Does Not Produce an Immediate Search Result

OpenAI’s crawler documentation gives an approximate 24-hour search-policy adjustment period. It does not promise a citation by that time.

Confirm delivery once, record the time, and avoid making several unrelated changes while evaluating the result. If access is working, review the page’s usefulness and relevance separately. Our guide to clear, source-supported content for AI search provides a related learning resource.

Someone Asks Whether the Same Rule Fixes Google AI Visibility

Google uses its own search controls. Its AI features documentation says supporting links in AI Overviews and AI Mode need indexed pages eligible for snippets; no special AI schema or text file is required. An OpenAI-specific group does not configure Googlebot.

When checking Google eligibility, use the appropriate indexing evidence. Our Google indexing guide explains that separate workflow.

What Should You Measure After the Change?

Maintain three separate records: technical access, observed answer citations, and identifiable human visits. They answer different questions.

EvidenceUseful interpretationWhat it does not establish
Delivered robots.txt and path checksThe intended crawl policy is available and applicableWhether the real crawler requested a particular page
A verified crawler request with useful content returnedThat request successfully retrieved the pageA citation, visitor session, or business conversion
A saved answer with a visible source linkThe page was cited in that observed answerVisibility across every question or future answer
Identifiable visitor sessions and key eventsMeasured website visits and recorded outcomesTotal mentions, unseen answers, or unidentifiable visits

For analytics interpretation, the current GA4 default channel definitions include AI Assistant. Its medium rule is ai-assistant, and a recognized assistant referrer can trigger that assignment. Google’s AI Overviews and AI Mode belong to Organic Search. Do not assume every assistant visit is categorized as Referral, or relabel unknown Direct visits as measured AI exposure.

Keep attribution work separate from access verification. Our GA4 beginner guide explains the broader measurement context. A crawl log belongs in the access record; a session belongs in the analytics record.

Final Configuration Checklist

  • The intended policy is recorded and matches the owner’s decision.
  • The active robots.txt source is identified.
  • The previous complete response has been saved for recovery.
  • Existing matching groups have been reconciled.
  • Required path restrictions remain in the relevant groups.
  • Cached responses have been refreshed where necessary.
  • The public file contains the intended plain-text rules.
  • Representative paths have been checked with the correct bot tokens.
  • Confirmed retrieval blocks have been reviewed at the responsible layer.
  • The change date, checks, and next review date are recorded.

Recheck this configuration after switching SEO plugins, changing hosts, modifying a CDN policy, or moving WordPress into a different directory. Keep the policy owner involved so a new default does not silently replace an intentional choice.

OAI-SearchBot and GPTBot checklist covering policy choices, exclusions, delivered responses, and security events.
Record intentional policy choices and verify the response WordPress serves.

Frequently Asked Questions

Can I use the sample if WordPress is installed in a subdirectory?

Adapt the paths before using it. The robots.txt response belongs at the host root, while the administrative directory may sit under the installation’s subdirectory. Identify the actual admin and AJAX paths, preserve intentional existing restrictions, and test a public article plus an excluded path before treating the configuration as complete.

Should I add a bot-specific group if the wildcard rules already work?

First check the effective policy. An explicit group can make a deliberate choice easier to maintain, but adding it changes which rules apply and can omit restrictions you intended to keep. Review the complete file, carry required exclusions into the group, and verify the same representative paths after the change.

How can I check the file without using a terminal?

Open the root robots.txt URL in a logged-out browser session and compare its text with your saved policy. Your browser’s developer tools can show the network response status and content type. If the response contains a challenge, login screen, or unexpected text, ask the hosting administrator to inspect delivery and caching.

What if my hosting panel does not show crawler logs?

Ask the host or security administrator for a narrow log sample covering the relevant page, timestamp, response status, request user agent, and source IP. Use the current official crawler reference to assess identity. If that evidence is unavailable, record the verification limit instead of claiming a genuine crawler successfully fetched the page.

Do I need a paid plugin, special schema, or llms.txt for this setup?

The task is to deliver the intended robots.txt policy and make the permitted public pages retrievable. The documented free editor or a maintained physical file can support that task. A special file, schema type, or paid plugin does not replace those checks, and this configuration does not promise inclusion in an answer.

Why might crawler requests increase without more GA4 sessions?

A crawler request and a visitor session measure different activity. Retrieving an article for processing does not mean a person clicked through to your site. Evaluate access through request evidence and human traffic through identifiable sessions. Keep citations and mentions in a separate observation record instead of treating crawl volume as audience growth.

Conclusion

Use the OAI-SearchBot vs GPTBot distinction to document a clear policy, then implement that policy in the source WordPress actually serves. Reconcile the applicable groups, preserve required restrictions, clear relevant caches, and verify public responses and representative paths.

Keep the resulting evidence specific: a working policy, a successful request, an observed citation, and a measured visit are different outcomes. That separation makes future troubleshooting easier and helps you evaluate changes without assuming results you have not observed.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top