AI Crawler Policy

AI Crawler Policy

AI Crawler Policy shows observed robots.txt status and selected known-agent rule parsing — separate from search, retrieval, and training concepts, without claiming crawler compliance.

Observed robots.txt and known-agent rules — separate from search, retrieval, and training, with no compliance guarantee.

WordPress only Limited — WordPress observed policy only

Agent request

Known crawler user-agent seen

Observed policy

robots.txt rule parsed

Allow / Block

Unclear rule

Ambiguous or missing directive

What goes wrong

Owners hear conflicting advice about AI crawlers but lack a plain view of what their public robots.txt actually says today.

What you can see now

Owners who need a plain view of what their public robots.txt says to known AI and search agents today, not a promise that crawlers will obey.

You see fetched robots.txt status, selected agent blocks, and honest separation of retrieval vs training concepts — then fix supported paths with approval.

How it works

AI Crawler Policy in practice

  1. Fetch

    Retrieve live robots.txt from the scanned origin.

  2. Parse

    Evaluate selected known-agent User-agent blocks.

  3. Separate

    Distinguish search, retrieval, and training concepts in the report.

  4. Report

    Show observed rules — not compliance certification.

  5. Act

    Apply supported WordPress fixes with preview/approval if a change is proposed.

Evidence inputs

What it reads

  • Fetched robots.txt from the scanned origin
  • User-agent blocks for selected known retrieval agents
  • Content-Signal or policy lines where present
  • Allow/Disallow paths relevant to public marketing routes
Example output

Sample based on the documented methodology

robots.txt: reachable, 200 OK

User-agent: GPTBot — Disallow: /private/ (observed)

Retrieval vs training: reported separately — no compliance guarantee

Boundaries

Control, limits, and availability

Availability
Limited — WordPress observed policy only
Read/write behaviour
Read-only. Agentic Bridge does not edit robots.txt unless a supported WordPress fix explicitly covers that file and you approve it.
Approval boundary
Policy findings describe what was observed. Changing robots.txt requires owner approval on supported paths only.

What we do not claim

  • WordPress only — Shopify does not ship robots.txt or known-agent parsing in the current audit engine.
  • Does not guarantee crawlers obey your robots.txt.
  • Does not observe crawler visit logs — policy text only.
  • Not a replacement for hosting or CDN firewall rules.

Security & privacy

  • Only public robots.txt is fetched — no authenticated admin surfaces.
  • Does not log or store visitor IP data for this check.
Evidence details for evaluators
Availability
Limited
Platform
WordPress only
Evidence source
WordPress live robots.txt fetch and known-agent parser
Last verified
24 Aug 2026
Read or write
Read
Requires approval
Yes if a supported robots.txt fix is proposed
Reversible
Yes where rollback exists for that change type
External provider
No
Plan availability
Included in public scan evidence
Related guides

Go deeper

Common questions

Questions about AI Crawler Policy

Does this block GPTBot automatically?+
No. It reports what your robots.txt says. Changes require owner approval on supported paths.
Is this the same as Discovery Control?+
No. Discovery Control governs Agentic Bridge outputs. Crawler Policy reads public robots.txt.

See what AI says about your business.

Run a free scan and find out in about two minutes. No card, no code, and nothing changes without your OK.

Free to scan · Works with WordPress & Shopify · You stay in control