The Patient Hoarder: On the Slow Accumulation of Discovery
There’s a certain rhythm, a predictable cadence, to the conventional wisdom of search engine discovery. We are told to optimize, to streamline, to remove friction. We build sitemaps that read like efficient train schedules and prune our sites of anything that might cause a crawler to linger too long or wander down a dark alley. The goal is to present a clean, well-lit storefront, where every product is on display and the path to purchase is a straight line. It’s a logic of scarcity, built around a concept we’ve come to fear: the crawl budget. We are frugal, careful not to waste a single robotic glance.
But what if we’ve misunderstood the nature of the collector? We imagine the web crawler as a harried commuter, rushing through a station, glancing at departure boards. I’d like to propose a different, more counterintuitive image: not the commuter, but the hoarder. Not the prospector panning frantically for gold, but the beachcomber who walks the same shore every day, slowly accumulating a collection of shells, sea glass, and oddities, whose value only becomes apparent in the aggregate, over years. This crawler is not in a hurry. Its economy is not one of scarcity, but of immense, almost incomprehensible abundance.
The Illusion of the Budget
The very term "crawl budget" suggests a limited resource that must be meticulously managed. This mindset leads us to hide our dusty archives, our experimental branches, our forgotten corners. We focus our efforts on the pristine, the new, the obviously valuable. But a crawler like Googlebot isn't a visitor with a tight schedule; it's a force of nature, a tide that returns again and again. Its "budget" is less a household allowance and more the budget of the ocean—vast, cyclical, and not something we can usefully ration. Our attempts to be overly helpful, to present only the most curated version of our domain, might actually be limiting the depth of understanding this patient entity can develop.
What if, instead of hiding our labyrinthine passages, we embraced them? A page that is crawled infrequently, over a period of years, isn't necessarily a wasted crawl. Each visit is a data point, contributing to a deeper, more nuanced map of the site's topology. A link discovered today, a piece of content recrawled six months from now—these are not inefficiencies. They are the slow, deliberate stitches in a vast tapestry. The crawler is building context, understanding relationships that aren't immediately apparent in a sitemap's sterile list of URLs. It learns not just what is present, but how the pieces of your world connect over time.
The most profound discovery isn't always the instantaneous one. It’s the connection made between a new piece of content and a three-year-old blog post that a persistent crawler, one that revisits the archives, finally weaves together. By trying to be hyper-efficient, we risk creating a shallow, two-dimensional site in the eyes of the index. We sacrifice depth and historical context for the illusion of optimal resource allocation. Perhaps, then, the best strategy isn't to frantically manage a budget, but to build a world rich enough, and interlinked enough, to reward a patient hoarder on its endless, meandering walk. Sometimes, the most valuable thing you can offer the crawl is not a shortcut, but a fascinating detour.
Notes & further reading
A few pages I came back to while writing this: