Skip to content
Oday Bakkour
Back to Knowledge Hub

Daily SEO Note — September 14, 2026: Cloudflare's AI Crawler Defaults Flip Tomorrow

Oday Bakkour profile photo
Oday Bakkour
12 min read
Share
Daily SEO Note — September 14, 2026: Cloudflare's AI Crawler Defaults Flip Tomorrow

Audit window: September 11 to 14, 2026, all times UTC. Monday extends the lookback to 72 hours to cover the weekend. Across the window the Google Search Status Dashboard reported no crawling, indexing, ranking, or serving incidents, and no broad core or spam update was announced. Everything below traces to a primary source with a date and a rollout status.

SEO for Content Writers

The most consequential editorial change this window is documentation, not an algorithm. Google updated its aggregator unit documentation on September 11, 2026, completing a set of pages that spell out, for the first time in one place, which Search experiences exist only in certain countries and who is allowed into them. If you publish for readers in the European Economic Area, the eligibility rules for hotels, flights, long-distance rail and bus, and products are now written down and citable.

Google Documents Its EEA Aggregator and Supplier Units

Google published a hub page on regional differences in Search experience (last updated September 8, 2026) alongside detailed pages for the aggregator unit (updated September 11, 2026) and the supplier unit (updated September 8, 2026). Rollout status: live in the EEA, documented rather than newly launched. The aggregator unit is a multi-provider block for Vertical Search Services such as online travel agencies, comparison shopping services, metasearch engines, and directories. The top-ranked provider is expanded by default and users can switch between the alternatives.

Who it affects: any publisher or business competing for EEA queries in four verticals only. Hotels, flights, long-distance trains and buses, and products. Outside those verticals and outside the EEA, nothing here applies to you. The split that matters editorially is between aggregators and direct suppliers. Aggregators must be approved as a Vertical Search Service and must feed Google data through direct feeds or real-time APIs, which is a commercial and engineering commitment, not something a content team can unlock alone.

Direct suppliers get the softer path. An individual hotel, airline, or a brick-and-mortar service provider such as a plumber qualifies for the supplier unit without submitting anything extra. Google's documentation is explicit that no data beyond what is accessible through web crawling is required, provided the site serves EEA users and is crawlable. The catch is that the supplier unit only surfaces alongside an aggregator unit, so it is a placement you inherit from the query shape rather than one you can target directly.

What to do differently in the next brief: if you run a localized EEA edition for any of these four verticals, treat the local-language page as a genuine article with real supply-side detail, not a translated shell. The supplier unit reads your crawlable page, so thin transcreation costs you the placement. What to stop doing: stop assuming a single English article with hreflang alternates covers EEA travel and product queries. For these verticals the regional experience is structurally different, and a duplicate does not compete.

What the Search Console Generative AI Report Actually Reports

The generative AI performance report finished its global rollout on August 31, 2026, after being announced in June 2026. It is worth restating what it measures, because the gap between what editors assume and what the Search Console Help documentation defines is where bad briefs come from. The report counts impressions, defined as how many times links to your site were shown to a user in a generative AI feature, broken out by page, country, device, and date.

Note what is absent from that list: query data. The report does not tell you which prompts or questions surfaced your page inside AI Overviews or AI Mode. Who it affects: every content team that has been asked for an AI visibility report. You can say which of your pages are being surfaced and in which countries. You cannot say what was asked to surface them, and no amount of pivoting the report will produce that.

Separately, and labeled unconfirmed: community reports circulating between September 10 and 13 attribute to Google's John Mueller an acknowledgement that Search Console assigns every link inside an AI Overview the position of the AI Overview block itself, rather than the position of the individual link. The Search Console Help page does not document position semantics for generative AI features either way, so this stays out of any conclusion you publish until Google documents it.

What to do differently in the next brief: build AI-surface reporting around pages and markets, and pair it with your own query research rather than implying the two are the same dataset. What to stop doing: stop promising query-level AI Overviews attribution to stakeholders. The data does not exist in the tool, and presenting inferred queries as measured ones is the kind of claim that does not survive scrutiny.

Tomorrow's Crawler Default Is an Editorial Decision, Not Just an Ops One

Section 2 covers the mechanics, but the choice belongs to whoever owns distribution. On September 15, 2026, Cloudflare changes its default settings for AI traffic (announced July 1, 2026; takes effect tomorrow). Cloudflare separates AI bots into three categories: Search, which indexes content to answer questions later and is expected to return referral traffic; Agent, which acts in real time on a person's behalf; and Training, which ingests content to train or fine-tune a model.

Under the new defaults, Training and Agent are blocked on pages that display ads, while Search remains allowed. This applies to new domains onboarding to Cloudflare. Every customer can opt out of the new defaults before September 15. Who it affects: publishers whose ad-supported articles are the same articles they want quoted in AI answers.

The editorial consequence is that Training and Agent are genuinely different bargains and deserve different answers. Blocking Training keeps your archive out of model weights but costs you nothing in AI citations, because citation comes through the Search category. Blocking Agent is the one that bites: agent fetchers are what retrieve your page when a reader asks an assistant about it right now, which is the closest thing to a referral that AI surfaces currently produce.

What to do differently: decide the Training question as a rights position and the Agent question as a distribution position, and write them down separately so the answer is not made by whoever last touched the CDN dashboard. What to stop doing: stop treating a single AI block-everything toggle as a content protection strategy. It is a distribution cut as much as a rights one.

Ranking and Policy Systems: Nothing Confirmed in the Window

No core update, spam update, Search Essentials revision, or Quality Rater Guidelines change was published between September 11 and 14, 2026. The Search Status Dashboard is clear and the Search Central documentation changelog logged no ranking-system entries in the window.

Apply to Your Next Brief

  • For EEA hotels, flights, rail and bus, or product articles, commission a real localized page with supply-side specifics, not a translation of the English original.
  • Confirm the localized page is crawlable and actually serves EEA users, because the supplier unit is built from crawled content and nothing else.
  • Reframe AI visibility reporting around pages, countries, and devices, and label it as impressions, never as queries.
  • Remove query-level AI Overviews attribution from any recurring report or client deck that currently promises it.
  • Get a written answer to the Training question and the Agent question separately before tomorrow, and give it to whoever owns the CDN.
  • Leave the block-level AI Overview position claim out of published analysis until Google documents position semantics.

SEO for Developers

One breaking change lands tomorrow and it is the only item here that cannot wait: Cloudflare's new AI traffic defaults take effect September 15, 2026, and they contain a coupling that can cost you Googlebot access if you opt into blocking Training without reading the note about multi-purpose crawlers. Everything else in the window is routine.

Breaking: Cloudflare AI Traffic Defaults Change on September 15, 2026

Rollout identifier and date: announced in the Cloudflare changelog as New options to manage AI traffic on July 1, 2026, effective September 15, 2026. Breaking for new domains onboarding to Cloudflare: Training and Agent crawlers are blocked by default on pages that display ads, while Search remains allowed by default. All customers can opt out of the new defaults before September 15.

The exact symptom if ignored is not the obvious one. Cloudflare's changelog states that multi-purpose crawlers that combine Search and Training will be affected by the new defaults to block Training. Googlebot, Applebot, and Bingbot are multi-purpose. If you enable Training blocking without carving them out, you are not trimming AI training access, you are dropping your primary search crawlers, and the failure is silent: no error in Search Console, just a decaying crawl rate followed by index loss over days.

The setting to change is in the Cloudflare dashboard under AI crawler controls, per zone. Do that before verifying anything in robots.txt, because CDN-level enforcement overrides robots.txt intent: robots.txt is a request that well-behaved crawlers honor under RFC 9309, while the CDN returns a block regardless. Treat the two as separate layers and make sure they agree.

public/robots.txt
# Policy layer: states intent. Enforcement happens at the CDN.
# Search + citation surfaces: allow.
User-agent: Googlebot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

# Training corpora: disallow.
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

Then verify that the CDN is not silently overriding that intent. The check that matters is whether Googlebot still gets a 200 after the defaults flip. Run it tomorrow, not today, because today's result proves nothing about tomorrow's defaults.

scripts/verify-crawler-access.sh
#!/usr/bin/env bash
# Run on 2026-09-15 after the Cloudflare defaults change.
set -euo pipefail
URL="https://example.com/"

for UA in \
  "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
  "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)" \
  "Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)"
do
  CODE=$(curl -s -o /dev/null -w '%{http_code}' -A "$UA" "$URL")
  echo "$CODE  $UA"
done
# Expect: 200 for Googlebot and OAI-SearchBot. GPTBot may be 403 by design.

Next.js 16.3.5 Adds CSP Nonces to Loading and Template Scripts

Version and date: Next.js v16.3.5, released September 11, 2026, as a stable patch backporting five fixes from canary. Non-breaking. The SEO-relevant entries are "Add CSP nonce to script tags of loading and template files", two next/image disk cache fixes, "Emit whole-app server NFTs when output: 'standalone' is used with an adapter", and "Fix use cache prerender signal retention".

The symptom if ignored applies only if you run a nonce-based Content Security Policy. Before this patch, the script tags Next.js emits for loading.tsx and template.tsx did not carry the nonce, so a strict policy blocked them. The visible failure is a route that renders its loading state and never resolves, which a crawler fetches as a permanently skeletal page. It looks like a rendering problem and is actually a header problem, which is why it tends to survive several rounds of debugging.

The setting to change is your upgrade pin, plus the middleware that generates the nonce and writes the CSP header. If you are not on a nonce-based CSP, the other four fixes still argue for taking the patch: the next/image disk cache changes prevent 0-byte and empty entries from poisoning the optimizer cache, which otherwise serves broken images into image results.

middleware.ts
import { NextResponse } from 'next/server'
import type { NextRequest } from 'next/server'

export function middleware(request: NextRequest) {
  const nonce = Buffer.from(crypto.randomUUID()).toString('base64')

  // With Next.js >= 16.3.5, loading.tsx and template.tsx script tags
  // receive this nonce, so 'strict-dynamic' no longer breaks them.
  const csp = [
    `default-src 'self'`,
    `script-src 'self' 'nonce-${nonce}' 'strict-dynamic'`,
    `style-src 'self' 'unsafe-inline'`,
    `img-src 'self' data: https:`,
  ].join('; ')

  const headers = new Headers(request.headers)
  headers.set('x-nonce', nonce)

  const response = NextResponse.next({ request: { headers } })
  response.headers.set('Content-Security-Policy', csp)
  return response
}

export const config = {
  // Keep crawlable static assets and metadata routes out of the matcher.
  matcher: [
    {
      source: '/((?!api|_next/static|_next/image|favicon.ico|sitemap.xml|robots.txt).*)',
      missing: [{ type: 'header', key: 'next-router-prefetch' }],
    },
  ],
}

Cloudflare WAF Ships Three Log-Only Rules the Same Day

Rollout identifier and date: the scheduled WAF release was announced September 8, 2026, for release on September 15, 2026. Non-breaking as shipped. Three rules are added with a default action of Log, not Block: SSRF - Cloud - 3 (a new detection), Version Control - Information Disclosure - Beta (merging into the existing Version Control - Information Disclosure rule), and Command Injection - Generic 10 (a new detection).

Nothing to change today, and the rules do not target crawlers. It earns a line here only because it lands on the same day as the AI traffic defaults, and WAF false positives are a classic cause of unexplained Googlebot 403s. If crawl errors appear on September 15, you now have two candidate causes in the same CDN and should rule out the crawler defaults first, since these rules only log.

The Data Path Behind EEA Aggregator and Supplier Units

Date: aggregator unit documentation updated September 11, 2026; supplier unit updated September 8, 2026. Informational, no action for most sites. The engineering split is sharp. Aggregator unit eligibility requires approval as a Vertical Search Service plus data delivered through direct feeds or real-time APIs, so it is an integration project with an approval gate in front of it, not a markup change.

The supplier unit requires the opposite: nothing. Google's documentation states that no data beyond what is accessible through web crawling is needed, as long as the site serves EEA users and is crawlable. For a direct supplier the entire engineering requirement is that Googlebot can reach and render the page from an EEA context. That makes geo-gating the thing to audit. If your edge redirects or challenges EEA traffic, or serves a consent wall that blocks rendering, you are ineligible for a placement you would otherwise get for free.

Dependency Advisories: Nothing New in the Window

No new advisories affecting SEO packages, sitemap generators, or crawler dependencies were published between September 11 and 14, 2026. Check the GitHub Advisory Database directly if you pin an older major of any of them.

Ship Today

  1. Open the Cloudflare AI crawler controls for every zone and record the current Search, Agent, and Training settings before the September 15 defaults land.
  2. If you block Training, explicitly allow Googlebot, Applebot, and Bingbot, because blocking Training blocks multi-purpose crawlers that also do Search.
  3. Update public/robots.txt so the stated policy matches the CDN enforcement, keeping Search and Training bots as separate decisions.
  4. Upgrade to Next.js 16.3.5, and if you run a nonce-based CSP, confirm loading.tsx and template.tsx routes resolve past their loading state.
  5. Schedule scripts/verify-crawler-access.sh for September 15 and alert on any non-200 for Googlebot.
  6. If you serve the EEA in travel or product verticals, audit geo-gating, consent walls, and edge redirects for crawlability from an EEA context.
Add Oday Bakkour as a preferred source on Google

Comments

Share your thoughts and join the conversation

Leave a Comment

Loading comments...
RELATED