FREE TOOL

Robots.txt Generator and Validator: Free Tool with AI Crawler Controls

This free robots.txt generator and validator has two modes. The Generator builds a robots.txt file visually: set rules for all bots, Googlebot, Bingbot, Yandex or Baiduspider, add your sitemap URL, and tick which AI crawlers to block. The Validator and Tester lets you paste an existing file, check the rules and test whether a specific URL is blocked.

Build or validate your robots.txt
Quick start
1Enter site settings
2Choose AI crawler rules
3Add rules
4Click Generate robots.txt

Site Settings

Blocking these keeps your content out of future model training. It does not affect whether AI chatbots can cite you today.
Blocking these removes your content from AI-generated answers right now. Most sites should leave these unchecked.

Rules

Your robots.txt

Upload to your site root as robots.txt

        

Paste Your robots.txt

Paste your robots.txt content above, then test a URL below.

Syntax Check

Direct answer: What is a robots.txt file?
A robots.txt file is a plain-text file at the root of a domain, such as example.com/robots.txt, that tells crawlers which paths they may or may not fetch. It follows the Robots Exclusion Protocol, standardised as RFC 9309 in 2022. It controls crawling, not indexing, and well-behaved bots follow it voluntarily.

What a Robots.txt File Does, in Plain English

Before a polite crawler fetches your pages, it asks for /robots.txt. The file answers one question: which parts of the site may I visit? You can keep crawlers out of admin areas, internal search results, cart and checkout paths and duplicate parameter URLs, which saves crawl effort for pages that matter. You can also tell crawlers where your sitemap is.

Robots.txt is a request, not a lock. Google, Bing and other major search engines obey it. Bad bots and scrapers do not. Never use it to hide private data.

Does Disallow Remove a Page from Google?

No, and this is the most common misunderstanding. Disallow stops crawling. It does not remove a URL from the index. If other pages link to a blocked URL, Google can still list it, often without a description. To keep a page out of search results, use a noindex meta tag or an X-Robots-Tag header and leave the page crawlable so Google can see the instruction. Block and noindex together defeat each other. You can generate the noindex tag with the Meta Tag Generator.

How to Use the Robots.txt Generator

  1. 1
    STEP 1 Enter site settings. Add your sitemap URL so it appears in the Sitemap line.
  2. 2
    STEP 2 Choose AI crawler rules. Tick the training bots you want to block, and leave search and citation bots alone unless you want to disappear from AI answers.
  3. 3
    STEP 3 Add rules. Pick a user-agent (all bots, Googlebot, Bingbot, Yandex or Baiduspider), choose Allow or Disallow, and enter a path. Use Add Rule for more.
  4. 4
    STEP 4 Click Generate robots.txt. Review the output.
  5. 5
    STEP 5 Upload and test. Save as robots.txt, upload to your domain root, then open yourdomain.com/robots.txt to confirm it loads.
Build or validate your robots.txt

Robots.txt Validator and Tester

Use the Validator and Tester tab when you already have a file. Paste it in, review the parsed rules and enter a URL to see whether a given user-agent is blocked. Run this after every site change, redesign, migration or CMS plugin update, because one stray line can block your whole site.

When you validate a robots.txt file, check for these problems:

  • A bare Disallow: / under User-agent: *, which blocks all crawling.
  • Rules that appear before any User-agent line.
  • Typos in directive names, such as Dissallow.
  • Paths without a leading slash.
  • Case mismatches in paths: /Blog/ and /blog/ are different.
  • A Sitemap line with a relative URL instead of a full one.
  • Blocked CSS or JavaScript files that Google needs to render pages.

Robots.txt Syntax Explained

DirectiveWhat it doesExample
User-agentNames the crawler the rules apply to. * means all.User-agent: Googlebot
DisallowBlocks a path from crawlingDisallow: /wp-admin/
AllowPermits a path inside a blocked areaAllow: /wp-admin/admin-ajax.php
SitemapGives the full URL of your sitemapSitemap: https://example.com/sitemap.xml
* and $Wildcard and end-of-URL match, supported by Google and BingDisallow: /*.pdf$

Rules that decide what gets blocked

  • Paths are case-sensitive. Disallow: /Private/ does not block /private/.
  • User-agent names are matched case-insensitively under the standard, but copy the exact token from the crawler's documentation anyway.
  • The most specific rule wins. In Google's implementation the longest matching path applies, and when Allow and Disallow tie, Allow wins.
  • A crawler follows only its own group. If a Googlebot group exists, Googlebot ignores the * group.
  • Crawl-delay is ignored by Google. Some other crawlers read it.
  • Size limit: Google reads the first 500 KiB of the file.

AI Crawlers: Training Bots vs Search Bots

Large AI companies run separate crawlers for different jobs. Training crawlers collect data for future models. Search or retrieval crawlers fetch pages to answer a live question. Blocking one does not block the other. This distinction is why the generator splits AI bots into two groups.

User-agentOwnerPurposeIf you block it
GPTBotOpenAITrainingOpt out of training only
OAI-SearchBotOpenAIChatGPT searchLose visibility in ChatGPT search
ChatGPT-UserOpenAIUser-triggered browsingLose visibility when users ask ChatGPT to open pages
ClaudeBotAnthropicTrainingOpt out of training only
Claude-SearchBot, Claude-UserAnthropicSearch and user requestsLose Claude citations
PerplexityBotPerplexitySearch indexLose Perplexity listings
Google-ExtendedGoogleGemini training and groundingNo effect on Google Search ranking
CCBotCommon CrawlOpen dataset used by many modelsOpt out of that dataset

A common stance is to block the training-only bots and allow the search bots, so you stop feeding future models while staying visible in AI answers. It is a business decision, not a technical one. Publishers who earn from AI citations usually allow search bots. Publishers worried about reuse may block both. Crawler names and purposes change, so check each provider's current documentation before publishing, including Anthropic's current bot names.

Robots.txt cannot force compliance. In August 2025, Cloudflare published a report accusing Perplexity of using undeclared crawlers that changed user agents and IP addresses to get around no-crawl directives. Perplexity disputed parts of the report. The takeaway is general: if you must keep a bot out, use firewall or server rules, not only robots.txt.

What to Block and What to Allow

  • Usually block: admin folders such as /wp-admin/, internal search results, login and cart paths, staging directories, and filter or session parameter URLs that create duplicates.
  • Never block: your sitemap, or CSS and JavaScript files that affect rendering. Google needs them to see the page as users do.
  • Always add: a Sitemap line with the full URL. The XML Sitemap Generator can build the file.
  • Handle with care: WordPress themes and plugins often add their own robots.txt rules. Check the file after changes.

Common Robots.txt Mistakes

  • Disallow: / on a live site. Often left over from staging. It blocks all crawling and can drop pages from results over time.
  • Using Disallow to deindex. It does not. Use noindex.
  • Blocking and noindexing the same page. Google cannot read the noindex if it cannot crawl.
  • Wrong case in paths. /Shop/ and /shop/ differ.
  • Blocking resources Google needs to render the page.
  • Copying a file from another site. Rules are specific to your URL structure.
  • Not testing after a migration.
  • Treating it as security. The file is public, and listing secret paths advertises them.

Who Uses a Robots.txt Generator

WordPress and Shopify store owners

Who need clean rules without hand-editing.

Developers

Moving a site from staging to production.

SEO consultants

Checking a client file during an audit.

Publishers

Deciding their position on AI crawlers.

Agencies

Producing a standard file for many small sites.

Honest Limits of This Tool

  • Robots.txt is voluntary. Compliant crawlers follow it, others may not.
  • The tester reports what the rules say. It cannot prove how Googlebot will behave on every URL. Use the URL Inspection tool in Google Search Console for Google's own view.
  • It does not remove indexed pages. Use noindex or the Removals tool for that.
  • It builds files for the common crawlers listed. Add unusual user-agents by hand.
  • AI crawler names and policies change. Review your file a few times a year.

Privacy

Robots.txt is a public file, so include nothing you want to keep private. Never list confidential directory names to hide them. The tool processes your input in the browser.

Open the tool

Frequently Asked Questions

What is a robots.txt file?

It is a plain-text file at your domain root that tells crawlers which URLs they may fetch. It follows the Robots Exclusion Protocol, standardised as RFC 9309. It controls crawling, not indexing.

Does Disallow remove a page from Google?

No. Disallow only stops crawling. A blocked URL can still appear in results if other pages link to it. Use a noindex tag or X-Robots-Tag header on a crawlable page to keep it out of search.

Is robots.txt case-sensitive?

Paths are case-sensitive, so /Blog/ and /blog/ are different. User-agent names are matched case-insensitively under the standard, but copy the exact token from the crawler's documentation.

How do I validate my robots.txt file?

Open the Validator and Tester tab, paste your file, review the parsed rules and test a URL against a chosen user-agent. Then confirm in Google Search Console, and fix any warnings before uploading.

Does blocking GPTBot stop ChatGPT from citing my site?

No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and user-triggered browsing uses ChatGPT-User. Blocking GPTBot only opts you out of training.

Does blocking Google-Extended affect my Google rankings?

No. Google-Extended controls whether content is used for Gemini training and grounding. It does not affect Googlebot, Search crawling or rankings.

Where do I put robots.txt?

Put it at the root of the host so it loads at https://yourdomain.com/robots.txt. Each subdomain needs its own file. A file in a subfolder is ignored.

Should I add my sitemap to robots.txt?

Yes. Add a Sitemap line with the full URL. Crawlers that read the file will find your sitemap without a manual submission.

Can robots.txt stop all AI bots?

It stops bots that choose to comply. Some crawlers ignore it. For firm control, add server or firewall rules, such as user-agent or IP blocks.

Is the generator free?

Yes. Both the generator and the validator are free to use.

Ready to try the Robots.txt Generator and Validator?

Next steps: open the Robots.txt Generator and Validator, build your file, then add the optional llms.txt file to point AI tools to your best pages.

open the Robots.txt Generator and Validator
For contributors
Got something worth sharing? Write for us.

Original, well researched guides are always welcome here.

  1. 1Read the guidelines
  2. 2Send us your pitch
  3. 3Our editors review it