Skip to content
Oday Bakkour
Back to Knowledge Hub

Daily SEO Note — September 19, 2026: Lighthouse Starts Auditing Your llms.txt

Oday Bakkour profile photo
Oday Bakkour
12 min read
Share
Daily SEO Note — September 19, 2026: Lighthouse Starts Auditing Your llms.txt

Audit window: 18 September 2026 06:00 UTC to 19 September 2026 06:00 UTC. Every item below traces to a primary source with a date and a rollout status. Community reports that Google has not confirmed are labelled as unconfirmed and are kept out of the action blocks.

1. SEO for Content Writers

The most consequential editorial change today is quiet and regional: Google updated its aggregator unit and supplier unit documentation to cover local business queries. The detail that matters for writers is that the supplier unit needs no feed and no structured data. Google populates it from the open web, which makes your on-page business copy the feed. Meanwhile, nothing on the Search Status Dashboard shows a September ranking update, so treat this week's volatility chatter accordingly.

Google's EEA Supplier Unit Now Answers Local Business Queries

On 18 September 2026 the Google Search Central documentation changelog logged that local business query support was added to the aggregator unit and the supplier unit. Rollout status: documented and live in the European Economic Area. Before today these units covered hotels, flights, long-distance trains and buses, and products. Local businesses are now in scope.

This affects one vertical rather than all content: sites that serve EEA users and are direct providers of a local service, such as a clinic, a restaurant, a plumber or a single hotel. Aggregators and comparison sites sit on the other side of the same feature, in the aggregator unit, and the supplier unit only appears when the aggregator unit does.

The concrete editorial instruction: on every location and service page aimed at an EEA market, put the facts a user would need to choose you in crawlable body copy rather than in an image, a widget or a script-rendered map. Service area, opening hours, the services you actually perform, and what makes you the direct provider rather than a reseller. Google's own documentation states you do not need to provide additional data beyond what is accessible through web crawling to appear in this feature, which means text on the page is the eligibility mechanism.

Stop treating local landing pages as thin doorways that exist only to hold a map embed and a phone number. That pattern was already weak; it now forfeits a surface that competitors with real copy can win without submitting anything.

No September Core or Spam Update Is Running

Rollout status as of 19 September 2026 06:00 UTC: no active ranking incident. The Google Search Status Dashboard ranking history lists the August 2026 spam update, which started on 18 August 2026 and ran for two days and sixteen hours, as the most recent confirmed event. There is no September core update and no September spam update.

This affects all content, and the correct action is no action. Writers are being told this week that a September update is reshuffling rankings. Google has not announced one. If your traffic moved, the cause is more likely to be a feature change on the results page, a seasonal demand shift, or a measurement artefact than an unannounced ranking system rewrite. Stop rewriting healthy articles in response to an update that has not been confirmed to exist, and stop citing third-party volatility trackers as if they were Google announcements.

Your llms.txt Stopped Being a Courtesy File

Lighthouse v13.5.0 shipped on 18 September 2026 at 16:18 UTC with a new audit group called agent discovery, which scores llms.txt alongside a new Agent Resource Discovery check. Rollout status: released, expected in PageSpeed Insights within roughly two weeks and in Chrome 156 DevTools. The engineering detail is in Section 2; the editorial consequence is that a file most teams treated as an optional gesture is about to appear as a pass or fail next to your performance score.

This affects all content, and it changes who owns the file. An llms.txt that lists every URL on the site is a sitemap with a different extension and it helps nobody. The file is an editorial selection: the pages you would hand a researcher who had time to read ten things, each with a one-line description in your own words explaining what question that page answers.

The instruction for your next brief is to decide, at outline stage, whether the finished piece earns a line in llms.txt. If it does, write that line as part of the brief rather than retrofitting it later. Pieces that earn it tend to share a shape: a definition that stands alone in its first sentence, a comparison the reader can lift as a table, or first-hand data nobody else holds. Stop publishing llms.txt as a generated dump of your sitemap.

The Search Profile Badge Turns Author Identity Into a Discover Lever

Google added documentation on 16 September 2026 for adding a Search profile badge to your website. Rollout status: documented, available to publishers and creators who have claimed a Search profile. A Search profile gathers a creator's content from across the web and social platforms into one destination on Google.

This affects publishers and bylined content rather than all pages, and it is the first time author identity has had a direct, documented distribution consequence: Google states that when readers follow your Search profile, the content linked from that profile becomes more likely to appear for that audience in Discover. Treat it as the same class of lever as preferred sources, not as a vanity badge. The editorial action is to claim profiles for the authors who actually carry the site's expertise, place the badge on author pages and high-traffic articles, and keep the linked social accounts current so the profile aggregates something worth following.

Google Discover Is Testing a 'Dive Deeper' AI Layer (Unconfirmed)

Reported on 18 September 2026 by Search Engine Roundtable: a Dive deeper option in Google Discover that opens an AI-generated topic summary with publisher links instead of sending the reader straight to an article. Rollout status: unconfirmed test. Google has published nothing about it, so this is detection only and it stays out of the checklist below.

If it ships, the unit of competition in Discover shifts from the headline to the paragraph a summary can quote. That is worth watching rather than planning around today. The one thing worth doing now costs nothing: make sure each Discover-oriented article states its central claim in a self-contained sentence early, because that sentence is what any summariser has to work with.

Apply to Your Next Brief

  • Add a crawlable-facts check to every EEA location and service page: service area, hours, services performed, and direct-provider status in body text, not in an image or a widget.
  • Decide at outline stage whether the piece earns a line in llms.txt, and write that one-line description into the brief itself.
  • Open every article with a self-contained sentence that states the central claim without needing the headline for context.
  • Claim Search profiles for your named authors and place the badge on author pages and your highest-traffic articles.
  • Do not rewrite healthy pages this week on the theory that a September core update is running. None is confirmed.
  • Stop shipping local landing pages whose only substance is a map embed and a phone number.

2. SEO for Developers

The most consequential engineering change today is Lighthouse v13.5.0, released 18 September 2026 at 16:18 UTC, which adds an Agent Resource Discovery audit and groups it with llms.txt under a new agent discovery category. Within about two weeks these audits reach PageSpeed Insights, which means sites that have never published an agent manifest will start seeing new failing audits on a report their stakeholders already read. Nothing in the window breaks an existing build, but two items below will change what your reports say.

Lighthouse v13.5.0 Adds Agent Resource Discovery and llms.txt Audits

Version and date: Lighthouse v13.5.0, published 18 September 2026 16:18 UTC. Non-breaking. The symptom if ignored is cosmetic but public: a new failing audit group on every PageSpeed Insights run once the release propagates, and the same failures inside Chrome 156 DevTools. The release also adds response header capture to the fetcher and tightens whitespace handling in llms.txt validation, so a previously tolerated malformed file may now be flagged.

The audit looks for an agent manifest at a well-known path. The underlying Agentic Resource Discovery specification was announced by Google in mid-2026 and defines a machine-readable index of the agentic resources a domain offers: MCP servers, agent interfaces, APIs and knowledge files. Lighthouse's own test fixture targets .well-known/ai-catalog. Note that spec revision v0.91 moved the canonical path to /.well-known/ard.json and kept /.well-known/ai-catalog.json as a path a client may also consult, so publish the file the audit looks for and confirm the current field list against the specification before you treat this example as final.

public/.well-known/ai-catalog.json
{
  "version": "0.91",
  "name": "Oday Bakkour",
  "description": "Fullstack development, SEO and AI engineering notes.",
  "resources": [
    {
      "type": "knowledge",
      "name": "Site documentation index",
      "url": "https://oday.dev/llms.txt",
      "mimeType": "text/markdown"
    }
  ]
}

Serve it as application/json. If you are on Next.js, a file in public/ is enough; the .well-known directory needs no special routing. Verify after deploy with a plain request rather than trusting the build, because some CDNs treat dot-prefixed directories as hidden and will return a 404 for a file that exists on disk.

Next.js Serves HTML-Limited Bots a Different Render Path Than Your Users

Rollout identifier: the 16.4 canary line, with the bots and crawlers section of the caching guide rewritten on 17 September 2026 (Next.js releases). Non-breaking as a release, but it documents a failure mode that is breaking when you hit it: a page that renders perfectly for a human can fail to render for a crawler.

The mechanism, from the Bots and crawlers section of the Streaming guide: Next.js detects HTML-limited bots by user agent and waits for generateMetadata to resolve before streaming page content, so metadata lands in the head. Under Cache Components, visitors and DOM-capable crawlers get the prerendered shell immediately, but HTML-limited bots skip that shell and re-render the page dynamically. Work that completed at build time now runs at request time. If your shell depends on build-time data or on anything unreachable in the request-time environment, the crawler gets an error while your monitoring, which requests as a browser, stays green.

The setting to know is htmlLimitedBots, introduced in 15.2.0 and typed as a RegExp. The trap is in the documentation: specifying it overrides the Next.js default list entirely, so a config written to add one internal crawler silently drops Googlebot's renderer, Bingbot, Twitterbot and Slackbot from blocking metadata. If you must set it, re-state the defaults you still need.

next.config.ts
import type { NextConfig } from 'next'

const config: NextConfig = {
  // WARNING: this REPLACES the Next.js default list, it does not extend it.
  // Re-state the default bots you still need blocking metadata for.
  htmlLimitedBots:
    /Mediapartners-Google|AdsBot-Google|Google-PageRenderer|Bingbot|Twitterbot|Slackbot|MyInternalCrawler/,
}

export default config

One related behaviour worth a regression test: once streaming has started the status code is already committed to 200, so a notFound() that fires mid-stream cannot become a 404. Next.js injects a robots noindex meta tag into the streamed HTML instead. Call notFound() before any await or Suspense boundary if you need a real 404, and add a monitor that fetches key routes with an HTML-limited bot user agent so you catch a shell that renders for people and breaks for crawlers.

Schema.org 30.1 Ships Retail and Digital Product Passport Vocabulary

Version and date: Schema.org release 30.1, 16 September 2026. Non-breaking; these are additions. The release adds a DigitalProductPassport class with supporting EnvironmentalProductDeclaration and DeclarationOfConformity types for EU digital product passports, retail feed properties including consumerNotice, isOftenBoughtWith and specification on Product, plus itemPopularity on Offer and minimumOrderValue on shipping settings. MedicalSpecialty gained Ophthalmology and Audiology.

The symptom if ignored is simply a missed opportunity rather than an error: new vocabulary is inert until a consumer supports it. Google has not announced rich result eligibility for any of these terms, so ship them for correctness and for the AI answer surfaces that read raw JSON-LD, not in expectation of a new SERP feature. The file to change is whichever module emits your product JSON-LD.

components/ProductJsonLd.tsx (rendered output)
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Product",
  "name": "Example Product",
  "sku": "EX-1024",
  "consumerNotice": "Contains a rechargeable lithium-ion battery.",
  "specification": "IP67 rated; 2400 mAh; USB-C PD 3.0",
  "isOftenBoughtWith": {
    "@type": "Product",
    "name": "Companion Charging Dock",
    "url": "https://example.com/products/charging-dock"
  },
  "offers": {
    "@type": "Offer",
    "price": "149.00",
    "priceCurrency": "EUR",
    "availability": "https://schema.org/InStock"
  }
}
</script>

Google Is Blocking Scrapers Harder: Verify Googlebot, Never Trust the User Agent

Reported 18 September 2026 by Search Engine Roundtable: Google appears to be more successful at blocking scrapers and tracking tools, with Bing reported to be showing human-verification interstitials as well. Rollout status: unconfirmed, no vendor statement. It stays out of the Ship today list, but it has a concrete engineering consequence today.

If third-party rank and SERP-feature data thins out or goes missing this week, treat it as a collection gap in your vendor's pipeline before you treat it as a ranking loss, and reconcile against Search Console rather than against a tracker. The second consequence runs the other way: as scraping gets harder, more traffic will arrive wearing a Googlebot user agent that is not Googlebot. User agent strings are free to forge, so any rule that grants crawler privileges, bypasses a paywall, lifts a rate limit or skips a bot challenge must verify by reverse DNS with a forward confirmation, per Verifying Googlebot and other Google crawlers.

scripts/verify-googlebot.sh
#!/usr/bin/env bash
# Reverse DNS + forward-confirm. A user agent string alone proves nothing.
set -euo pipefail
IP="${1:?usage: verify-googlebot.sh <ip>}"

HOST="$(dig -x "$IP" +short | sed 's/\.$//')"
[ -n "$HOST" ] || { echo "no PTR record for $IP -> NOT Googlebot"; exit 1; }

case "$HOST" in
  *.googlebot.com|*.google.com|*.googleusercontent.com)
    if dig +short "$HOST" | grep -qx "$IP"; then
      echo "verified Googlebot: $HOST ($IP)"
    else
      echo "forward lookup mismatch -> NOT Googlebot: $HOST"; exit 1
    fi
    ;;
  *) echo "hostname not in a Google domain -> NOT Googlebot: $HOST"; exit 1 ;;
esac

No New Security Advisories for SEO Packages in This Window

Checked against the GitHub Advisory Database for 18 to 19 September 2026: nothing new for next-seo, next-sitemap, the Nuxt or Astro sitemap integrations, or schema tooling. The most recent SEO-adjacent advisories remain the two June 2026 high-severity ultimate-sitemap-parser issues, GHSA-p5wc-9w9r-m232 (XML entity expansion) and GHSA-8823-qg2x-pv9f (gzip decompression bomb), plus CVE-2026-46497, an SSRF via sitemap-derived URLs in crawlee. If you parse third-party sitemaps anywhere in your pipeline, those three are still the ones to check your lockfile against.

Ship Today

  1. Publish an agent manifest at public/.well-known/ai-catalog.json and confirm it returns 200 with application/json from the edge, not just from the build output.
  2. Validate your llms.txt against Lighthouse v13.5.0 before the audits reach PageSpeed Insights, and fix whitespace issues the tightened validator now flags.
  3. Audit next.config.ts for an htmlLimitedBots override and re-state every default bot you still need, since the option replaces the default list rather than extending it.
  4. Add a synthetic check that fetches your top ten routes with an HTML-limited bot user agent and asserts a complete document, so a shell that renders for people and fails for crawlers is caught in CI.
  5. Move any notFound() call ahead of the first await or Suspense boundary on routes where a real 404 status matters.
  6. Replace every user-agent-only Googlebot rule at the CDN or WAF with reverse DNS plus forward confirmation.
  7. Add the Schema.org 30.1 retail properties to your product JSON-LD emitter where you already hold the data.
Add Oday Bakkour as a preferred source on Google

Comments

Share your thoughts and join the conversation

Leave a Comment

Loading comments...
RELATED