The Librarian's Unwritten Volume: On the Myth of Total Indexing
There exists a quiet, persistent dream in the corners of our work: the fantasy of the complete index. It’s the notion that with enough diligence, the right tools, and a perfectly structured sitemap, we can present our entire library—every single page, every footnote, every scrap of content—to the search engine, and that it will, in turn, faithfully catalog it all. We speak of ‘crawl budget’ as if it were a mere accounting problem, a ledger we can balance to ensure every entry is counted. This is our industry’s most seductive and damaging fiction.
We operate as if the crawler is an obedient archivist, waiting patiently for our instructions. We furnish the sitemap, that table of contents we so meticulously crafted, believing it to be a binding contract. But a crawler is not a librarian; it is a forager. It is driven by a complex, ever-shifting calculus of perceived value, user signals, and its own operational constraints. It doesn’t seek to build a perfect replica of our site in its index. It seeks to build a useful one for its users.
This is the core of the misunderstanding. The goal of a search engine is not preservation, but discovery. Its index is not an archive but a tool for answering questions. From this perspective, the crawl budget isn’t a quota to be filled, but a measure of attention to be earned. Pages that are seldom linked, rarely updated, or provide little unique value are not ‘misfiled’ by the crawler; they are consciously deprioritized. The engine makes a judgment call, deciding that its resources are better spent elsewhere. It is, in its own algorithmic way, practicing a form of curation.
To rage against this is to rage against the nature of the medium itself. The web is not, and has never been, a place where everything is meant to be found. It is a dynamic, living ecosystem where obscurity is the default state. Our job, then, shifts from one of exhaustive accounting to one of compelling advocacy. We are not filing clerks stamping documents for processing; we are advocates for our content, building a case for why it deserves a crawler’s precious attention.
This is a more profound and human challenge than simply optimizing a robots.txt file. It asks us to consider not just how a page is built, but why it should exist in the vastness of the web. What question does it answer? What curiosity does it satisfy? What connection does it make? The pages that get found are not just the ones that are technically accessible; they are the ones that whisper a compelling reason to be sought. The rest, much like the unwritten volumes in a great library, remain in the quiet dark, known only to the keeper of the shelves.
Notes & further reading
A few pages I came back to while writing this:
- Gilbert, AZ
- The Mason's Mortar: On the Strength of a Single Internal Link
- Peoria, AZ
- The Miller's Full Grain: On the Myth of the Sparse Field
- Surprise, AZ
- The Geographer's Final Silence: On the Necessary Omission in Every Index
- Elk Grove, CA
- Pasadena, CA
- New Haven, CT
- Stamford, CT
- Washington, DC
- one area's overview
- a practical rundown