The Museum Archivist's Whisper: On Curating for the Future Crawler

Last week, I spent an afternoon in a small maritime museum. It wasn't the polished exhibits that stuck with me, but a brief glimpse into the back room. An archivist was meticulously logging a crate of newly donated sailor's journals, her movements slow and deliberate. She wasn't just storing them; she was creating a context—dating entries, cross-referencing ship names, noting the type of ink used. This wasn't for the visitor tomorrow, but for the researcher in fifty years. It struck me that her work is a profound, almost spiritual parallel to how we should think about building websites for crawlers we can't yet imagine.

We often discuss crawl budget and sitemaps as technical logistics, a kind of inventory management. But the archivist's craft reframes it as an act of intentional legacy. Her primary concern is not immediate discovery, but enduring, meaningful access. She knows that simply having the artifact is useless if its story and connections are lost. In our world, a page is just a digital artifact. Without the proper context—clear signals of its place in the site's structure, its relationship to other pages, its temporal relevance—it becomes a ghost in the archive, present but perpetually unfindable.

Preservation Through Structure

An archivist would never toss a priceless letter into a random drawer. It gets a catalogue number, is placed in a climate-controlled box within a specific collection, and its metadata is woven into a finding aid. On the web, our HTML is that finding aid. A clean, semantic hierarchy of headings isn't just for readability; it's the archival box and the catalogue number. A logical, shallow URL structure is the clearly labeled shelf in the stacks. A sitemap is not just a ping to a search engine; it's the master index donated to the central library, a statement of what exists and where it can be found.

The archivist also practices ruthless, thoughtful selection. Not every scrap of paper is kept; space and attention are finite. This is the most direct lesson for our crawl budget. The archivist's 'crawl budget' is her lifetime and her storage. She allocates it to items of enduring value, provenance, and completeness. When we let thin, duplicate, or outdated content linger without purpose, we are forcing the future crawler to sift through our digital attic junk, wasting its limited time and attention on artifacts that tell no worthwhile story.

Ultimately, the archivist works for a patron she will never meet. She trusts that her system will hold, that her labels will be understood, that her curatorial decisions will make sense to a future mind. When we build a site, we are, in a way, whispering to a future crawler—an agent of discovery whose algorithms and priorities will inevitably change. By borrowing the archivist's mindset, we build not for the spike of traffic, but for the long, slow discovery. We curate a collection that is navigable, meaningful, and built to last, so that when the crawler of tomorrow finally arrives, it doesn't just find pages. It understands a library.

Notes & further reading

A few pages I came back to while writing this: