Skip to main content
TechSEO Vitals

Cloudflare Is About to Block Googlebot for Sites That Blocked AI Training

The 'Block AI bots' toggle you flipped last year starts blocking Googlebot too. Cloudflare now judges crawlers on every purpose – and the strictest rule wins.

Share on

Last year, blocking AI bots felt like the obvious move. Cloudflare made it one click. A lot of people pressed it.

In a few weeks, that click starts blocking Googlebot.

Not a warning. Not a preference in a text file. Cloudflare blocks the request before it can read your robots.txt file.

What actually changes

Cloudflare now sorts bots by behavior. Search, Agent, Training.

Two things change in a couple of weeks. The one getting covered is a new default. New domains get Training and Agent blocked on pages with ads. Only new domains. If your site is already on Cloudflare, that change never reaches you.

The other one is why you're reading this.

Crawlers get judged on all of their behaviors, not one. The most restrictive rule wins.

Googlebot crawls for Search. It also crawls for Training. Same with Applebot and BingBot.

So if Training is blocked, Googlebot is blocked too. Cloudflare names all those three bots in the announcement.

That includes the legacy "Block AI bots" toggle. You don't need to have touched the new controls. The old button counts.

You pressed it last year. You lose Googlebot in a few weeks.

Crawling drops first. Then indexing. Then rankings. Search Console fills with errors that point at your server. Not at a dashboard you haven't opened since last summer.

Why they're doing it

Almost every project I run sits behind Cloudflare. It's still the best infrastructure platform I know. One bad default won't change that.

And their reasoning holds up.

One crawler does search and training at the same time. So you can't say no to training without losing search. That was never a fair choice.

Cloudflare decided to stop letting training hide behind search. If Google won't split the crawler, treat it as both.

But I go back and forth on this one.

Separate crawlers is the right principle. If you're taking my content for three purposes, I should be able to say yes to one and no to another.

Google's side isn't stupid either. They already fetched the page. Fetching it again for a different purpose means more requests and more load on my server, for content they already have.

One crawl, many uses, is lighter on your site. Anyone who has watched AI crawlers hammer their server knows why that matters.

Both sides have a point. Splitting the crawler gives you control and costs you bandwidth. Nobody has proposed the version that fixes both.

One thing is certain. Cloudflare picked a side. The pressure is aimed at Google. The blast radius includes you.

What to do before September 15

Open your Cloudflare security settings. Check whether Training is blocked, through the new AI traffic controls or the old toggle.

If it is, and you want Googlebot to keep crawling, opt out. Cloudflare has a setting for exactly this. It tells them to leave the multi-purpose crawlers alone. The option is open until the deadline.

Then the bigger question. What does blocking training actually buy you?

If your content is the product, blocking training is a real business decision. But you still need the opt-out. Blocking training is one thing. Losing Google is another. Until September, you could do the first without the second.

Here's the setup I'd run. Opt out of the multi-purpose rule so Googlebot keeps crawling. Then block the training crawlers by name instead. Robots.txt for the ones that comply, an edge rule for the ones that don't.

Single-purpose bots blocked by name. Multi-purpose bots left alone. You keep Google and you keep training out.

But be honest about whether any of this is you. For most sites it isn't a real decision at all. A SaaS, a distributor, a shop. Nobody was ever going to pay to read your content. Blocking training just removes you from the next model.

The line that does nothing

If managed robots.txt is on, Cloudflare wrote a block into your file that you didn't write. Legal-sounding text, then a line like Content-Signal: search=yes,ai-train=no.

It looks like control. But it isn't.

Mueller was asked about it on Reddit recently. As far as he knows, no crawler uses it. No LLM either. It has no effect at all. It just adds bloat to your file.

Cloudflare doesn't really argue. Their own docs call content signals a preference, "rather than issuing blocks directly."

If you're allowing training, you can leave the lines alone. They're not hurting anything. They're just not doing anything.

If you're blocking it, don't count on them. Block the crawlers by name instead.

I want this to work

The idea isn't wrong. That's what makes it frustrating.

Robots.txt was built to say where a crawler can go. It was never built to say what happens to your content afterward. That gap is real, and it gets wider every month.

The fix isn't complicated on paper. Cloudflare, Google, Microsoft and OpenAI in a room, agreeing on one set of rules. Written by the people who build the crawlers. Followed because they wrote it.

That's how robots.txt happened in the first place. Nobody enforced it. The people who mattered just agreed to use it.

I don't see that meeting happening. The incentives point the other way. Google has no reason to help you say no to training. Cloudflare has no reason to wait for Google.

So my bet is it ends in law instead. Not syntax one company invented. A real rule, with a regulator behind it, that crawlers follow because they have to. Europe moved first on the copyright side. This feels like the same road.

Cloudflare is early, not wrong.

But someone has to write it for humans. Right now it's a comma-separated line in a file most site owners have never opened. Under a paragraph of legal text.

I asked a few friends who own websites what ai-train=no does for them. Nobody knew. One didn't know the line was in their file at all.

A standard that needs a technical SEO to translate it isn't finished.

Look at the two settings side by side. The one that sounds like a policy does nothing. The one that sounds like a small technical detail can pull Googlebot off your site.

One of them got a launch post, a legal framework, and a generator tool. The other got a paragraph.

Go check your settings.

Martin Stepanek

Martin Stepanek

Enterprise Technical SEO Consultant

I am a developer-led enterprise technical SEO consultant. 10+ years building the web before I started fixing it, and I still ship production code today. I read the responses your site actually sends, tell you which findings are worth a sprint, and take the fix to your developers myself.

Every two weeks

The technical SEO newsletter engineers and SEOs actually finish

One specific problem, taken apart, with a position on what to do about it — argued against the primary documentation and against what I see in audits. Then three stories from the last two weeks, picked and explained.

Mersudin ForbesMersudin ForbesMark Williams-CookMark Williams-CookAleyda SolisAleyda Solis
Recommended by industry leaders

Subscribe

A new episode every two weeks. Unsubscribe in one click.

By subscribing, I agree to the Privacy Policy and Terms and Conditions.

No spam, ever. Unsubscribe at any time.
Technical SEO notes