Skip to content
Oday Bakkour
Back to Knowledge Hub

Daily SEO Note — August 10, 2026: Google's Own Guidance Says llms.txt Does Nothing for Search

Oday Bakkour profile photo
Oday Bakkour
11 min read
Share
Daily SEO Note — August 10, 2026: Google's Own Guidance Says llms.txt Does Nothing for Search

1. SEO for Content Writers

The most consequential editorial fact this Monday is an absence. Across the full 72-hour weekend window Google shipped nothing to its ranking, spam, or policy surfaces, so the item that should change your briefs is not news at all — it is the guidance already sitting on Google's own documentation, which contradicts a large share of the "optimize for AI" advice currently in circulation. Google states plainly that it ignores llms.txt, that structured data is not required for generative AI features, and that you do not need to rewrite or chunk anything for machines. If your briefs contain those instructions, they are working from a premise Google has publicly rejected.

Google's Ranking, Spam, and Policy Surfaces Logged Nothing All Weekend

No verified change. The Search Status Dashboard reports no incidents across Crawling, Indexing, Ranking, and Serving, last updated 9 August 2026 at 23:06 PDT — 10 August 06:06 UTC. There is no open core update, spam update, or Discover incident.

The Search Central Blog has published nothing in August 2026; its newest posts are from July. The documentation changelog's most recent entry is dated 29 July 2026, covering Search Console analysis of social and video platform content. Rollout status across all three surfaces: nothing in flight.

Google Says llms.txt Does Nothing and AI Answers Need No Special Format

What changed: nothing this weekend — but Google's guide to optimizing for generative AI features (last updated 10 July 2026, live documentation) is far more restrictive than most published GEO advice, and it is worth treating as the operative brief. On llms.txt, the guidance is one sentence: Google Search ignores them. They do not affect rankings or AI-feature eligibility.

Who it affects: all content, and especially any team that has been restructuring articles on the theory that AI systems need shorter, machine-shaped text. Google says structured data is not required for generative AI search and there is no special schema.org markup to add; that there is no requirement to break content into tiny pieces; and that you do not need to write in a specific way just for generative AI. It also warns that seeking inauthentic mentions across the web is less helpful than it appears.

What to do differently: put the effort back into unique, first-hand perspective — original testing, proprietary data, direct experience — and into supporting images and video, which the guide names explicitly. Write for a human reader with clear sections and descriptive headings. What to stop: stop maintaining llms.txt as a Google Search tactic, stop fragmenting articles into answer-sized chunks, and stop commissioning generic listicles that recycle what is already indexed. Note the separate case for other engines is unchanged — this is a statement about Google Search only, and OpenAI's crawlers remain a different decision.

The Search Console Generative-AI Control Is a Single Site-Wide Switch

What changed: still rolling out, no status change this weekend. The Search generative AI control lives at Settings > Search generative AI in Search Console and decides whether your content can appear in AI Overviews, AI Mode, and generative AI features in Discover. Google is rolling it out to a subset of website owners.

Who it affects: every editorial team whose site has more than one person with Search Console access. It is binary and site-wide — include (the default), exclude, or inherit from a parent property. Google states it is not used as a ranking or inclusion signal affecting other parts of Search, so flipping it does not touch your blue-link rankings; exclusions typically take effect within one to two days.

What to do differently: before your next publishing cycle, confirm in writing who owns this toggle. A single exclude setting removes an entire traffic surface without producing any ranking symptom an editor would recognise — impressions in AI features simply stop while organic positions look normal. What to stop: stop assuming AI-surface visibility is purely earned. On this site it is also a permission, and permissions get changed by people who are not reading your content strategy.

Generative AI Performance Reports Give You Impressions — and Only Impressions

What changed: no status change in the window; the reports were announced on the Search Central Blog in June 2026 and are still rolling out to a subset of websites so Google can test and gather feedback before wider release.

Who it affects: anyone reporting on AI-surface performance. The generative AI performance report covers AI Overviews and AI Mode, and what it counts is impressions — how many times links to your site were shown inside a generative AI feature. You can group by page, country, and date, and the newest data is preliminary, marked with a dotted line, and may still move.

What to do differently: report AI-surface visibility as an impression trend and name it as such in the document. If you need behaviour rather than exposure, you still need analytics on the landing side. What to stop: stop writing CTR or average-position narratives for AI Overviews and AI Mode from this report — those columns are not what the report measures, and a chart that implies otherwise will be wrong in a way nobody catches for a quarter.

Unconfirmed: Early-August Volatility Chatter Still Has No Primary Source

Labelled unconfirmed and deliberately excluded from the checklist below. Community trackers and forum threads reported sharp ranking movement in the 1–6 August period, with inconsistent and sometimes contradictory symptoms — positions improving while sessions fell, or traffic dropping with no visible rank change.

Google has confirmed no update, and the Search Status Dashboard logged no ranking incident for that period. Treat it as unattributed until a primary source says otherwise; do not rewrite pages or brief a recovery project against it. If you saw a real drop, the useful next step is separating Search from Discover and from analytics collection before assuming an algorithm cause.

Apply to Your Next Brief

  • Delete any llms.txt requirement from briefs written for Google Search visibility — Google states it ignores the file.
  • Remove instructions to chunk articles into short machine-readable answer blocks; Google says there is no such requirement.
  • Replace "add schema for AI" line items with a first-hand evidence requirement: original testing, proprietary data, or direct experience.
  • Commission supporting images or video as part of the brief, not as an afterthought — Google's guidance names them explicitly.
  • Name the owner of the Search Console generative-AI toggle in your publishing runbook this week.
  • Label AI Overviews and AI Mode numbers as impressions in every report; do not present them as clicks, CTR, or position.
  • Drop mention-farming and inauthentic-citation tactics from outreach briefs.
  • Do not brief any recovery work against the unconfirmed early-August volatility.

2. SEO for Developers

Nothing SEO-relevant shipped to the frameworks, CDNs, or advisory feeds inside the 72-hour window. The most consequential engineering item is therefore a dated deadline rather than a release: on 15 September 2026 Cloudflare changes its default AI-bot policy, and the setting that decides your outcome is one you have to visit before then. Everything else below is a verification task — confirming that what you believe you serve is what crawlers actually receive.

Scheduled: Cloudflare Flips AI-Bot Defaults on 15 September 2026

Identifier and date: Cloudflare changelog post 2026-07-01-ai-traffic-options, published 1 July 2026, with the companion announcement blog. Rollout status: scheduled, effective 15 September 2026. Breaking for new domains, non-breaking for existing zones — but only if you leave them alone.

Cloudflare splits AI traffic into three categories: Search (crawlers that index so they can answer questions later), Agent (automated activity acting in real time for a person), and Training (crawlers taking content to train or fine-tune a model). From 15 September, new domains onboarding to Cloudflare get Training and Agent blocked on pages that display ads, while Search stays allowed. Existing customers keep their current settings unless they change them.

The symptom if ignored: multi-purpose crawlers are matched on all of their behaviours, so a zone that blocks Training also blocks Googlebot, Applebot, and BingBot. That is an indexing outage dressed as a bot policy. The setting is at Security > Settings > Configure AI bot traffic policies, and Cloudflare allows opting out of the new defaults at any point before 15 September. Decide the three categories independently at the origin as well, so the intent survives a CDN change.

public/robots.txt
# Three independent decisions, stated explicitly.
# Search: allow — this is how you get cited and referred.
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

# Training: deny — no model training on this content.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Bytespider
Disallow: /

# Agent / user-triggered: allow reads, keep them off internals.
User-agent: ChatGPT-User
User-agent: Claude-User
Disallow: /admin/
Disallow: /api/
Allow: /

Sitemap: https://example.com/sitemap.xml

Verify the Edge Is Not Quietly Overriding robots.txt Intent

Non-breaking, but it is the check that makes the item above meaningful. Under RFC 9309, robots.txt is advisory: it states intent and well-behaved crawlers honour it. A WAF or bot rule enforces something different — it returns 403 or a challenge regardless of what the file says. When those two disagree, the file is fiction.

The symptom if ignored: your robots.txt allows OAI-SearchBot, your edge rule blocks it under a broad "AI bots" category, and you lose citation eligibility while every audit tool reports the policy as correct. Test from outside your network, per user agent, and compare the status codes against what you intended.

scripts/check-crawler-access.sh
#!/usr/bin/env bash
# Confirm the edge honours robots.txt intent, per crawler.
set -euo pipefail
ORIGIN="${1:-https://example.com}"

declare -A AGENTS=(
  [OAI-SearchBot]="Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)"
  [GPTBot]="Mozilla/5.0 AppleWebKit/537.36 (compatible; GPTBot/1.1; +https://openai.com/gptbot)"
  [PerplexityBot]="Mozilla/5.0 AppleWebKit/537.36 (compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)"
  [Googlebot]="Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
)

for name in "${!AGENTS[@]}"; do
  for path in /robots.txt /; do
    code=$(curl -s -o /dev/null -w '%{http_code}' \
      -A "${AGENTS[$name]}" --max-time 20 "${ORIGIN}${path}")
    printf '%-16s %-12s %s\n' "$name" "$path" "$code"
  done
done

# 200 on both = reachable. 403/503 on / while robots.txt allows it
# means an edge rule is overriding your stated intent.

The Weekend Advisory Stream Carried No SEO Dependency

No action, but worth recording so your monitor does not re-alert. Reviewing the GitHub Advisory Database for 9 and 10 August 2026, roughly twenty-five advisories were published and none affect a CMS platform, SEO plugin, sitemap or crawler library, or web framework. All were unreviewed, and the stream was dominated by MCP server packages, alongside GStreamer, and assorted IoT and desktop utilities.

The symptom if ignored is the inverse of the usual one: an advisory monitor tuned to fire on any new GHSA will page you all weekend for MCP wrappers that are not in your dependency tree. Constrain the query to the ecosystems you actually ship, and re-check that your filter still matches the packages you depend on.

scripts/seo-dep-advisories.sh
#!/usr/bin/env bash
# Only alert on advisories in ecosystems this site actually ships.
set -euo pipefail
SINCE="${1:-2026-08-07}"

gh api -X GET /advisories \
  -f ecosystem=npm \
  -f published=">=${SINCE}" \
  -f per_page=100 \
  --jq '.[]
        | select(.summary != null)
        | {ghsa: .ghsa_id, sev: .severity,
           pkg: [.vulnerabilities[].package.name] | unique | join(","),
           summary: .summary}
        | select(.pkg | test("sitemap|seo|robots|crawler|next|nuxt|astro|svelte"; "i"))'

# Empty output for 2026-08-09 and 2026-08-10 is the expected result.

Lighthouse Is Still on 13.4.1 — Pin the Version You Audit Against

Version and date: Lighthouse v13.4.1, released 20 July 2026. Rollout status: shipped to npm; the release notes state it is expected in Chrome 152 DevTools and in PageSpeed Insights within two weeks of that date. No August release exists, so 13.4.1 remains current as of this window.

Non-breaking, but it changes what your audits report. The release adds detailed error messages when fetching robots.txt or llms.txt fails — useful, because the long-standing failure mode was a bare "unable to download a robots.txt file" that gave no cause. It also enables the Agentic Browsing category via the PSI API. Note that Google's own Search guidance says llms.txt is ignored by Search, so treat a failing llms.txt audit as an agent-readiness signal, not an SEO defect.

The symptom if ignored: local Lighthouse and PageSpeed Insights disagree, and you spend an afternoon chasing a robots.txt audit difference that is only a version skew. Read the version out of the API response rather than assuming. We attempted this check live while writing and the keyless endpoint returned HTTP 429, so use a PSI API key for anything you intend to run on a schedule.

scripts/psi-lighthouse-version.sh
#!/usr/bin/env bash
# Read the Lighthouse version PSI is actually running before
# comparing its output against a local run.
set -euo pipefail
URL="${1:-https://example.com}"
: "${PSI_API_KEY:?set PSI_API_KEY - the keyless endpoint returns 429}"

curl -sS --max-time 90 -G \
  'https://www.googleapis.com/pagespeedonline/v5/runPagespeed' \
  --data-urlencode "url=${URL}" \
  --data-urlencode "key=${PSI_API_KEY}" \
  --data-urlencode 'strategy=mobile' \
| python3 -c 'import json,sys; r=json.load(sys.stdin)["lighthouseResult"]; \
print("lighthouse:", r["lighthouseVersion"]); \
print("fetched:  ", r["fetchTime"]); \
print("categories:", ", ".join(r.get("categories", {})))'

# Compare against: npx lighthouse --version

Quiet Everywhere Else, With Dates

Recorded so the absence is verifiable rather than assumed. Next.js stopped at v16.3.1-canary.9, published 8 August 2026 at 23:44 UTC, with nothing released on 9 or 10 August; the stable line's last tag in the window was v15.5.23 on 7 August. Nothing in either touches the Metadata API, app/sitemap.ts, app/robots.ts, or revalidation semantics.

The Cloudflare developer changelog has no entries dated 8, 9, or 10 August; its most recent is 7 August. The Vercel changelog published nothing touching redirects, rewrites, middleware, ISR, Cache-Control, or image optimization in the window. Schema.org remains at release 30.0 dated 19 March 2026 — no vocabulary change. For context on where the CDN-side conversation is heading, Cloudflare's agent-readiness and answer-engine work landed 6 August, just outside this window, and is worth reading before the September deadline.

Ship Today

  1. Open Security > Settings > Configure AI bot traffic policies in Cloudflare and record your current Training, Agent, and Search selection before 15 September 2026.
  2. Confirm that blocking Training in that panel will not also block Googlebot, Applebot, and BingBot on your zone; opt out of the new defaults if it would.
  3. Run the crawler-access check against production and reconcile any status code that disagrees with public/robots.txt.
  4. Split robots.txt into explicit Search, Training, and Agent groups so the intent is legible without reading a dashboard.
  5. Constrain the advisory monitor to the ecosystems you ship, so MCP-package noise stops paging the on-call.
  6. Pin the Lighthouse version in CI and read lighthouseVersion from the PSI response before filing any audit discrepancy.
  7. Confirm nobody flipped the Search Console generative-AI control to exclude, and write down who owns it.
Add Oday Bakkour as a preferred source on Google

Comments

Share your thoughts and join the conversation

Leave a Comment

Loading comments...
RELATED