The Index at the End of the World: On the Futility of Hoarding Crawl Budget
There’s a recurring dream in our line of work. It goes something like this: if only we had more, if only the crawler would stay longer, if only the budget were infinite, then our entire domain would be rendered in perfect, searchable clarity. Every last page, from the flagship product announcement to the forgotten blog post from a decade ago, would be bathed in the light of the index. It’s a fantasy of completeness, a librarian’s ambition to catalogue every single leaf in the forest. And I’m here to suggest it’s not only a fantasy, but a fundamentally misguided one.
The common wisdom, the ‘received idea’ I want to question, is the very premise that a larger crawl budget is an unalloyed good. We talk about it as if it were a natural resource, a barrel of oil or a cistern of water to be jealously guarded and maximized. We optimize our server response times, we sculpt our internal links, all to eke out a few more seconds of the crawler’s precious attention. This is what I call the ‘hoarder’s mindset.’ We want to stockpile this abstract currency, believing that the sheer quantity of crawled pages translates directly into value.
But what is the value of an indexed page that nobody will ever seek? What is the purpose of ensuring a crawler can access the 500th page of an image gallery or the sprawling, unread archive of meeting minutes from 2015? The hoarder believes that the act of inclusion is itself the victory. Yet, the index is not a trophy case for our digital ephemera; it’s a tool for connection. By fighting for the inclusion of everything, we devalue the currency of discovery for the pages that truly matter.
Consider the crawler’s journey not as a frantic scavenger hunt for every last byte, but as a guided tour for a very important, but incredibly busy, visitor. You wouldn’t lead a guest through every closet, utility room, and dusty attic in your house. You’d show them the spaces where life is lived, where conversations happen, where value is created. The goal isn’t to prove the sheer scale of your domain’s real estate, but to present its most meaningful contours.
The hard truth is that most websites are littered with pages that are, for all practical purposes, islands. They have no inbound links from other sites, they generate no external signals of interest, and they serve no ongoing purpose to a living audience. Hoarding crawl budget to ensure these pages are indexed is like building a lighthouse on a cliff that no ship’s captain has ever seen or will ever need to see. The light shines, but it illuminates nothing but its own solitude.
Perhaps a better metaphor is that of a gardener, not a hoarder. The gardener doesn’t cherish every single leaf. They prune. They deadhead. They understand that by removing the dying and the redundant, they direct the plant’s energy—and the sunlight—towards the most vibrant blooms. In our digital gardens, a finite crawl budget is not a limitation to be resented, but a focusing mechanism to be embraced. It forces us to ask the difficult question: is this page worthy of the light? If the answer isn’t a resounding yes, then perhaps the most powerful optimization is to let it fade gracefully into the background, making room for what truly deserves to be found.
Notes & further reading
A few pages I came back to while writing this:
- Torrance, CA
- The Quiet Protocol: How to Ask for Less Indexing, and Mean It
- Aurora, CO
- Don't Feed the Beast: Why Obsessing Over Crawl Budget Misses the Point
- Colorado Springs, CO
- The Weaver of Cobwebs: On How a Spider Brought Us the Searchable Web
- Denver, CO
- Fort Collins, CO
- Lakewood, CO
- Thornton, CO
- Bridgeport, CT
- Hartford, CT
- New Haven, CT