Our first self-improving SEO loop ran every five minutes. It rebuilt pages, crawled the live site, checked canonical rules, generated reports, and attempted to pull several external signals. The machine was busy enough to be visible in Cloudflare analytics.

The problem wasn’t that crawling had no value. The problem was frequency and purpose. We were using a production proof step as a heartbeat, even when nothing had been deployed since the previous run.

What the crawl was good at

Those are valuable checks after a deployment. They don’t need to happen twelve times an hour when the release is unchanged.

  • Finding 404s, broken assets, bad redirects, noindex mistakes, and canonical mismatches.
  • Confirming that the live edge served the release we intended to publish.
  • Checking sitemap and robots files after a release.
  • Creating a reproducible deployment proof bundle.

What the crawl couldn’t create

The redesigned loop runs cheap local checks frequently. Full live crawls happen after a release or on a limited schedule. Evidence pulls follow the pace of search data, and the expensive step runs only when there’s something new to prove.

  • A stronger business entity or better local prominence.
  • Reviews, citations, trustworthy external references, or real customer proof.
  • A useful page that answers the query better than established competitors.
  • Search demand, click-through rate, or conversion intent.

The schedule now follows change

Source files, metadata, internal links, schema, and sitemap membership can be checked locally whenever code changes. A live crawl belongs after a deployment, during a scheduled health check, or when an external signal suggests something broke. Search and rank measurements belong on a slower cadence because Google doesn’t meaningfully re-evaluate a page every five minutes.

This makes the system quieter and more useful. Each run has a named purpose, a bounded cost, and a report that distinguishes a new observation from an old artifact being regenerated again.