Robots.txt Tester Widget

Embed a robots.txt tester in your technical SEO guide. Readers paste a robots.txt file, a URL and a crawler name, and see whether the URL is allowed or blocked and exactly which line decides it.

SEO & Content Tool Runs in your browser Free · no ads

Customize your widget

Theme
Auto follows the visitor's light/dark setting.
Style
Attribution on your page
Optional and entirely your choice. The exact line is shown in the code below; it links to the tool with rel="nofollow".
More options
Starting values
Leave blank to use the widget's defaults. Visitors can still change every value.

Live preview

Exactly what your visitors will see

Embed code

<iframe src="https://a2z.tools/embed/w/robots-txt-tester" title="Robots.txt Tester by A2Z Tools" width="100%" height="760" style="border:0;width:100%" loading="lazy" allow="clipboard-write"></iframe>

A plain iframe. Works everywhere, including site builders that strip scripts. Adjust height if your content needs more room.

Works with

How it works

The widget implements the matching rules of RFC 9309, the Robots Exclusion Protocol standard that Google and Bing follow. A crawler obeys the group whose User-agent matches its name (case-insensitive), merging groups with the same name; only if none matches does it fall back to the * group - it never combines its own group with *. Within the group, the rule with the longest matching path wins, and if an Allow and a Disallow match with the same length, Allow wins. * matches any characters and a $ at the end anchors the end of the URL. Paths are compared after normalising percent-encoding, and /robots.txt itself is always allowed. The table lists every rule in the group and whether it matched.

Method

  • Group: exact product-token match (case-insensitive), else the * group, else allow everything
  • Winner: longest matching rule path; tie -> Allow (RFC 9309 section 2.2.2)
  • Wildcards: * = any sequence; trailing $ = end of URL; empty Disallow matches nothing
  • Normalisation: non-ASCII -> UTF-8 %XX; %XX of unreserved characters decoded; hex upper-cased

Limitations

  • Crawl-delay, Sitemap, Host and Clean-param lines are accepted but play no part in the allow/block answer; Google ignores Crawl-delay, while Bing honours it.
  • It cannot know how a particular crawler deviates from RFC 9309, nor whether a bot obeys robots.txt at all.
  • Only the pasted text is tested: redirects, HTTP 4xx/5xx responses for robots.txt and the 500 KiB size limit crawlers apply are not simulated.
  • Meta robots tags and X-Robots-Tag headers, which control indexing rather than crawling, are outside its scope.

Where publishers use it

  • Technical SEO courses explaining crawl directives
  • Site-migration checklists where staging rules must not reach production
  • Documentation for AI-crawler and bot policies
  • Developer blogs debugging why a page is not being crawled

Questions

Does Disallow remove a page from Google?

No. It only asks crawlers not to fetch the page; it can still be indexed from links. Use a noindex tag or header on a crawlable page to keep it out of results.

Why doesn't the * group apply to Googlebot here?

When a group names the crawler, the crawler follows only that group and ignores *. Repeat any shared rules inside the named group.

How are Allow and Disallow conflicts resolved?

The longest matching path wins; on an exact tie, Allow wins.

Does the widget fetch my live robots.txt?

No. Paste the file's contents; the A2Z Robots.txt Checker can fetch a live site.

Sources

  1. RFC 9309: Robots Exclusion Protocol - IETF . Group selection, longest-match and Allow-on-tie rules, * and $ wildcards, percent-encoding; September 2022
  2. How Google interprets the robots.txt specification - Google Search Central . Google's handling of groups and unsupported fields

Cite or recommend this tool

If you reference this tool in an article, course or documentation, these formats are ready to copy. They are optional - nothing is added to your site unless you paste it.

A2Z Tools Robots.txt Tester
https://a2z.tools/robots-txt-checker

Preview