Robots.txt Generator and Validator: Free Tool with AI Crawler Controls
This free robots.txt generator and validator has two modes. The Generator builds a robots.txt file visually: set rules for all bots, Googlebot, Bingbot, Yandex or Baiduspider, add your sitemap URL, and tick which AI crawlers to block. The Validator and Tester lets you paste an existing file, check the rules and test whether a specific URL is blocked.
Build or validate your robots.txtSite Settings
Rules
Your robots.txt
robots.txt
Paste Your robots.txt
Syntax Check
A robots.txt file is a plain-text file at the root of a domain, such as example.com/robots.txt, that tells crawlers which paths they may or may not fetch. It follows the Robots Exclusion Protocol, standardised as RFC 9309 in 2022. It controls crawling, not indexing, and well-behaved bots follow it voluntarily.
What a Robots.txt File Does, in Plain English
Before a polite crawler fetches your pages, it asks for /robots.txt. The file answers one question: which parts of the site may I visit? You can keep crawlers out of admin areas, internal search results, cart and checkout paths and duplicate parameter URLs, which saves crawl effort for pages that matter. You can also tell crawlers where your sitemap is.
Robots.txt is a request, not a lock. Google, Bing and other major search engines obey it. Bad bots and scrapers do not. Never use it to hide private data.
Does Disallow Remove a Page from Google?
No, and this is the most common misunderstanding. Disallow stops crawling. It does not remove a URL from the index. If other pages link to a blocked URL, Google can still list it, often without a description. To keep a page out of search results, use a noindex meta tag or an X-Robots-Tag header and leave the page crawlable so Google can see the instruction. Block and noindex together defeat each other. You can generate the noindex tag with the Meta Tag Generator.
How to Use the Robots.txt Generator
-
1STEP 1 Enter site settings. Add your sitemap URL so it appears in the Sitemap line.
-
2STEP 2 Choose AI crawler rules. Tick the training bots you want to block, and leave search and citation bots alone unless you want to disappear from AI answers.
-
3STEP 3 Add rules. Pick a user-agent (all bots, Googlebot, Bingbot, Yandex or Baiduspider), choose Allow or Disallow, and enter a path. Use Add Rule for more.
-
4STEP 4 Click Generate robots.txt. Review the output.
-
5STEP 5 Upload and test. Save as robots.txt, upload to your domain root, then open yourdomain.com/robots.txt to confirm it loads.
Robots.txt Validator and Tester
Use the Validator and Tester tab when you already have a file. Paste it in, review the parsed rules and enter a URL to see whether a given user-agent is blocked. Run this after every site change, redesign, migration or CMS plugin update, because one stray line can block your whole site.
When you validate a robots.txt file, check for these problems:
- A bare Disallow: / under User-agent: *, which blocks all crawling.
- Rules that appear before any User-agent line.
- Typos in directive names, such as Dissallow.
- Paths without a leading slash.
- Case mismatches in paths: /Blog/ and /blog/ are different.
- A Sitemap line with a relative URL instead of a full one.
- Blocked CSS or JavaScript files that Google needs to render pages.
Robots.txt Syntax Explained
| Directive | What it does | Example |
|---|---|---|
| User-agent | Names the crawler the rules apply to. * means all. | User-agent: Googlebot |
| Disallow | Blocks a path from crawling | Disallow: /wp-admin/ |
| Allow | Permits a path inside a blocked area | Allow: /wp-admin/admin-ajax.php |
| Sitemap | Gives the full URL of your sitemap | Sitemap: https://example.com/sitemap.xml |
| * and $ | Wildcard and end-of-URL match, supported by Google and Bing | Disallow: /*.pdf$ |
- Paths are case-sensitive. Disallow: /Private/ does not block /private/.
- User-agent names are matched case-insensitively under the standard, but copy the exact token from the crawler's documentation anyway.
- The most specific rule wins. In Google's implementation the longest matching path applies, and when Allow and Disallow tie, Allow wins.
- A crawler follows only its own group. If a Googlebot group exists, Googlebot ignores the * group.
- Crawl-delay is ignored by Google. Some other crawlers read it.
- Size limit: Google reads the first 500 KiB of the file.
AI Crawlers: Training Bots vs Search Bots
Large AI companies run separate crawlers for different jobs. Training crawlers collect data for future models. Search or retrieval crawlers fetch pages to answer a live question. Blocking one does not block the other. This distinction is why the generator splits AI bots into two groups.
| User-agent | Owner | Purpose | If you block it |
|---|---|---|---|
| GPTBot | OpenAI | Training | Opt out of training only |
| OAI-SearchBot | OpenAI | ChatGPT search | Lose visibility in ChatGPT search |
| ChatGPT-User | OpenAI | User-triggered browsing | Lose visibility when users ask ChatGPT to open pages |
| ClaudeBot | Anthropic | Training | Opt out of training only |
| Claude-SearchBot, Claude-User | Anthropic | Search and user requests | Lose Claude citations |
| PerplexityBot | Perplexity | Search index | Lose Perplexity listings |
| Google-Extended | Gemini training and grounding | No effect on Google Search ranking | |
| CCBot | Common Crawl | Open dataset used by many models | Opt out of that dataset |
A common stance is to block the training-only bots and allow the search bots, so you stop feeding future models while staying visible in AI answers. It is a business decision, not a technical one. Publishers who earn from AI citations usually allow search bots. Publishers worried about reuse may block both. Crawler names and purposes change, so check each provider's current documentation before publishing, including Anthropic's current bot names.
Robots.txt cannot force compliance. In August 2025, Cloudflare published a report accusing Perplexity of using undeclared crawlers that changed user agents and IP addresses to get around no-crawl directives. Perplexity disputed parts of the report. The takeaway is general: if you must keep a bot out, use firewall or server rules, not only robots.txt.
What to Block and What to Allow
- Usually block: admin folders such as /wp-admin/, internal search results, login and cart paths, staging directories, and filter or session parameter URLs that create duplicates.
- Never block: your sitemap, or CSS and JavaScript files that affect rendering. Google needs them to see the page as users do.
- Always add: a Sitemap line with the full URL. The XML Sitemap Generator can build the file.
- Handle with care: WordPress themes and plugins often add their own robots.txt rules. Check the file after changes.
Common Robots.txt Mistakes
- Disallow: / on a live site. Often left over from staging. It blocks all crawling and can drop pages from results over time.
- Using Disallow to deindex. It does not. Use noindex.
- Blocking and noindexing the same page. Google cannot read the noindex if it cannot crawl.
- Wrong case in paths. /Shop/ and /shop/ differ.
- Blocking resources Google needs to render the page.
- Copying a file from another site. Rules are specific to your URL structure.
- Not testing after a migration.
- Treating it as security. The file is public, and listing secret paths advertises them.
Who Uses a Robots.txt Generator
WordPress and Shopify store owners
Who need clean rules without hand-editing.
Developers
Moving a site from staging to production.
SEO consultants
Checking a client file during an audit.
Publishers
Deciding their position on AI crawlers.
Agencies
Producing a standard file for many small sites.
Honest Limits of This Tool
- Robots.txt is voluntary. Compliant crawlers follow it, others may not.
- The tester reports what the rules say. It cannot prove how Googlebot will behave on every URL. Use the URL Inspection tool in Google Search Console for Google's own view.
- It does not remove indexed pages. Use noindex or the Removals tool for that.
- It builds files for the common crawlers listed. Add unusual user-agents by hand.
- AI crawler names and policies change. Review your file a few times a year.
Privacy
Robots.txt is a public file, so include nothing you want to keep private. Never list confidential directory names to hide them. The tool processes your input in the browser.
Frequently Asked Questions
What is a robots.txt file?
It is a plain-text file at your domain root that tells crawlers which URLs they may fetch. It follows the Robots Exclusion Protocol, standardised as RFC 9309. It controls crawling, not indexing.
Does Disallow remove a page from Google?
No. Disallow only stops crawling. A blocked URL can still appear in results if other pages link to it. Use a noindex tag or X-Robots-Tag header on a crawlable page to keep it out of search.
Is robots.txt case-sensitive?
Paths are case-sensitive, so /Blog/ and /blog/ are different. User-agent names are matched case-insensitively under the standard, but copy the exact token from the crawler's documentation.
How do I validate my robots.txt file?
Open the Validator and Tester tab, paste your file, review the parsed rules and test a URL against a chosen user-agent. Then confirm in Google Search Console, and fix any warnings before uploading.
Does blocking GPTBot stop ChatGPT from citing my site?
No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and user-triggered browsing uses ChatGPT-User. Blocking GPTBot only opts you out of training.
Does blocking Google-Extended affect my Google rankings?
No. Google-Extended controls whether content is used for Gemini training and grounding. It does not affect Googlebot, Search crawling or rankings.
Where do I put robots.txt?
Put it at the root of the host so it loads at https://yourdomain.com/robots.txt. Each subdomain needs its own file. A file in a subfolder is ignored.
Should I add my sitemap to robots.txt?
Yes. Add a Sitemap line with the full URL. Crawlers that read the file will find your sitemap without a manual submission.
Can robots.txt stop all AI bots?
It stops bots that choose to comply. Some crawlers ignore it. For firm control, add server or firewall rules, such as user-agent or IP blocks.
Is the generator free?
Yes. Both the generator and the validator are free to use.
Related Free Tools
Schema Markup Generator
Create JSON-LD for 16 page types and check which still earn rich results.
Open the toolAEO Readiness Checker
Score your content against 15 AI citation signals.
Open the toolE-E-A-T Score Checker
Check 20 on-page trust signals and get a prioritised fix list.
Open the toolGoogle SERP Simulator
Preview your title and description in Google on desktop and mobile.
Open the toolReady to try the Robots.txt Generator and Validator?
Next steps: open the Robots.txt Generator and Validator, build your file, then add the optional llms.txt file to point AI tools to your best pages.
open the Robots.txt Generator and Validator