Should You Block AI Crawlers?

No, you should not block AI crawlers if you run a store, because blocking bots like OAI-SearchBot pulls your products out of the answers ChatGPT, Claude, and Gemini give shoppers. Blocking makes sense for some publishers who sell attention. It almost never makes sense for a merchant who sells products.

AI Crawlers Do Three Different Jobs

AI crawlers do three different jobs on your site, because training a model, building a search index, and fetching a page for a live chat are separate tasks run by separate bots. Training crawlers collect content to teach future models. Search crawlers build the index that AI answers cite, and user-triggered fetchers grab a page live because someone just asked about it.

OpenAI documents its bots along exactly those lines. GPTBot trains models, OAI-SearchBot powers ChatGPT search, and ChatGPT-User fetches pages when a user asks. Anthropic runs the same split with ClaudeBot, Claude-SearchBot, and Claude-User.

Each one costs you something different when blocked. Block GPTBot and your content is left out of future training. Block OAI-SearchBot and your store drops out of ChatGPT search answers, and block ChatGPT-User and a live product lookup fails right when a shopper asked about you.

So lumping them together is where most blocking mistakes start. A store owner reads a headline about AI training, copies a block list from a news site, and locks out the search bot in the same edit.

Blocking the Search Bot Pulls Your Store Off the Shelf

Blocking the search bot pulls your store off the shelf, because OAI-SearchBot is what puts your pages in ChatGPT search answers, and that is where shoppers now ask what to buy. Blocking GPTBot alone only opts you out of training data, and OpenAI says nothing else changes. The damage comes when you block the whole family.

Plenty of sites have made that trade. Cloudflare’s crawl data shows 14% of the top 10,000 domains block AI crawlers in robots.txt, with GPTBot the most blocked bot of all. Meanwhile GPTBot request volume grew 305% in a year and ChatGPT-User grew 2,825%, which tells you how fast shoppers are actually using these tools.

Picture what that trade looks like on a Tuesday night. A shopper asks ChatGPT which heated jacket to buy for a Chicago winter. OAI-SearchBot has already indexed your competitor’s product page, and yours returned a blocked response, so the answer names three jackets, none of them yours, and the shopper never learns you exist.

That is why I never block a search or user-fetch bot to solve a training objection. If you sell products, the citation is the prize, and how those citations get chosen is what I break down in how ChatGPT picks its sources.

Publishers Have a Reason to Block, Stores Do Not

Publishers have a reason to block AI crawlers and stores do not, because a publisher’s product is the content itself. When ChatGPT summarizes a recipe or a news story, the reader never visits, the ad never loads, and the publisher earns nothing. For them, blocking training bots is a real negotiating lever, and some have licensing money on the line.

Your store is the opposite case. An AI answer cannot summarize away a heated jacket or a mattress base. It can only describe one and point somewhere, and you want the pointing to land on you.

That pointing is worth more than it sounds. When an assistant names your product, that is shelf placement you did not pay for, in front of a buyer who asked to be sold to. No ad, no bidding, no coupon, just your product named as the answer.

The buyers it sends are good ones too. I cover the conversion data in my post on whether AI citations convert into customers, and the short version is that AI referrals out-convert almost everything else. The publisher playbook is built for a business model you do not have, so stop borrowing it.

Keep Robots.txt Permissive and Leave Googlebot Alone

Keep your robots.txt permissive and leave Googlebot alone, because every bot attached to an assistant that recommends products can send you a buyer. For Shopify clients I allow OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, and PerplexityBot, and I leave GPTBot and ClaudeBot alone too, since training data is how future models learn your brand exists. I only block a specific bot that hammers the server or belongs to a scraper with no audience.

Just know the first step is reading what is already there. Open yourdomain.com/robots.txt in a browser and search for GPTBot, OAI-SearchBot, and ClaudeBot. On Shopify the default file is fine, and the trouble usually comes from a robots.txt.liquid edit someone made a year ago and forgot about.

Also, the mechanics are simple. A user-agent line names the bot, a disallow line names what it cannot touch, and Google’s robots.txt documentation covers the syntax. Robots.txt is voluntary, so reputable bots obey it while scrapers ignore it, which means it manages relationships more than it enforces anything.

The one trap is Google. Google-Extended only controls Gemini training and grounding, while AI Overviews run on regular Googlebot, so blocking Googlebot to escape AI Overviews would delete you from Google entirely. I will concede the control layer is messy right now, and I wrote up whether you need an llms.txt file separately, since that proposal comes up in the same breath.

Getting Crawled Is Step One, Getting Picked Is the Game

Getting crawled is step one and getting picked is the game, because an open robots.txt only puts your pages in the pool, and the assistant still chooses which store to name. Hundreds of millions of people use ChatGPT every week. A permissive file just means you are eligible when one of them asks about your category.

From there, the work looks like regular ecommerce SEO. Clear product titles, direct answers on collection pages, and reviews that say what the product is for all give the search bot something worth citing. For ChatGPT specifically, a product feed matters too, and I cover that setup in getting Shopify products into ChatGPT Shopping.

The problem is that none of that work counts if the bot cannot get in. Fix access first, then spend the rest of the quarter on the pages themselves.

Want Your Products in the AI Aisle?

Blocking AI crawlers protects publishers who sell attention. You sell products, and AI assistants are becoming the aisle where shoppers ask for recommendations, so your job is the opposite: stay crawlable and get cited.

If you want to know where you stand, book a call and I will read your robots.txt with you and tell you what I would change first. Want it built for you instead? Take a look at my AI search services.

Similar Posts